EPAM Systems
Senior Site Reliability Engineer
Warsaw, PL · Posted 1w ago
Job Description
We are seeking a Senior Site Reliability Engineer to own the reliability, observability, and operational health of production AI systems, bridging the gap between deployment and long-term operability while embedding cost, security, and quality discipline into every solution's lifecycle.ResponsibilitiesOwn deployment end-to-end, including infrastructure as code, CI/CD pipelines, and environment management on Azure, ensuring every environment is rebuildable from sourceBuild LLM-aware observability with traces on every model call, production quality signals such as eval sampling, drift detection, and guardrail-trigger rates, plus cost and latency dashboardsDefine and defend SLOs covering availability, latency, and quality objectives per solution, balancing delivery speed against stability with data-driven error budgetsRun incident management, including on-call models, runbooks written before incidents occur, and blameless postmortems afterwardManage the cost of intelligence by monitoring token economics per solution, wiring in budgets and alerts, and conducting capacity planning proactivelyKeep the security posture current through patching, secret rotation, access reviews, and audit readiness across the solution's entire lifecycleShape operability requirements before handover, ensuring they land in the pod's definition of done, and run hypercare jointly with sign-off on what will be operatedFeed operational patterns, failure modes, and cost learnings back to the pods and the ArchitectRequirements5+ years of experience operating cloud production systems, with a track record in scaling, defining SLOs, managing on-call rotations, and automating manual work awayExpertise in Azure IaaS/PaaS operations, infrastructure as code (Terraform/Bicep), and CI/CD toolingProficiency in observability stacks with LLM tracing, container orchestration, and Python/Bash automationKnowledge of FinOps basics for AI workloadsFamiliarity with daily AI use in operations work, including incident triage, runbook drafting, log analysis, and automation codeWe offerWe gather like-minded people:Top tech minds driving innovation in AI, cloud and digital platform modernizationSupportive team and agile, startup-like cultureHybrid by design mode and opportunity to work remotely within PolandChance to work abroad for up to 60 days annuallyBusiness-driven relocation opportunitiesWe provide growth opportunities:Career development programsThought leadership, mentoring, soft skills and well-being programsCertification (Anthropic, Gemini, GCP, Azure, AWS)English classesWe cover it all:Stable payParticipation in the Employee Stock Purchase Plan with a 15% discountBenefits package (health insurance, multisport, shopping vouchers)Referral bonuses up to $2,000Offices featuring entertainment and relaxation zones, table tennis and football, free snacks, coffee and moreCorporate, social and well-being eventsPlease, note:Benefits listed above are available to employees onlyWe are open for working with Contractors. Terms of B2B cooperation agreements are agreed individuallyWe will reach out to selected candidates exclusivelyEPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments.