DCG
Senior DevOps Engineer (AI & Platform Operations)
Warszawa, PL · Posted 3w ago
Job Description
As a recruitment company, DCG understands that every business is powered by experienced professionals. Our management style and partnership approach enable us to meet your needs and provide continuous support. Due to our ongoing growth and the large number of recruitment projects we undertake for our partners, we are currently looking for:Senior DevOps Engineer (AI & Platform Operations)Responsibilities:Incident & Problem Management: Own the RCA process for production incidents — diagnose, resolve, and put preventive measures in place so issues don't recurProduction Monitoring & Support: Continuously monitor service health, detect anomalies early, and act before they become incidentsDeployment Execution: Trigger and oversee release deployments through existing CI/CD pipelines; troubleshoot failed deployments and coordinate rollbacks when neededEnvironment Oversight: Keep Pre-Production and Production environments stable and aligned — not building them from scratch, but ensuring they behave as expected day to dayRunbook & Knowledge Management: Document operational procedures, known issues, and resolution steps to build a reliable knowledge base for the teamCross-team Collaboration: Work shoulder-to-shoulder with development and platform teams to triage issues, clarify operational requirements, and close the feedback loop between prod and devIdentify recurring pain points and propose automation or tooling to reduce toilImprove observability coverage — dashboards, alerts, log queries — to catch issues fasterContribute to service continuity initiatives and disaster recovery drills Requirements:5+ years in IT operations, application support (2nd/3rd line), or a similar production-facing roleProven track record of owning incidents end-to-end — from alert to RCA to prevention2+ years working within an ITIL framework (incident, problem, change management)Experience working in Agile delivery environments alongside development teamsExcellent English communication skills — able to explain technical issues clearly to both engineers and non-technical stakeholdersProficiency with log analysis and alerting tools: Splunk, Apica, SysdigObservability tooling: Prometheus, Grafana — reading dashboards, tuning alertsComfortable operating services running on Kubernetes (checking pod health, reading logs, triggering restarts — not cluster administration)Familiarity with Jenkins pipelines to execute and troubleshoot deploymentsRelational databases (Oracle, DB2) — querying, interpreting execution plans, identifying data-related incidentsWorking knowledge of Spring/Hibernate application behavior, Kafka message flows, XML/JSON payloads — enough to trace an issue through the stackNice to have:Java/J2EE development background (helps enormously when reading stack traces and working with dev teams)IBM Datastage operational experienceScripting (Bash, Python) for automation of repetitive operational tasksAnsible for applying configuration changes in controlled operational scenarios Offer:Private medical careCo-financing for the sports cardConstant support of dedicated consultantEmployee referral program