RITS

Python Developer (SRE)

Remote · Posted yesterday

salary not listedcontract

Job Description

We are looking for Python Developer (SRE). Rate: 160-180 PLN/h net + VAT (B2B)100% remote Project DescriptionJoin a global technology organization focused on ensuring the reliability, stability, and operational excellence of large-scale production systems. The role sits at the intersection of Incident Operations, Site Reliability Engineering (SRE), and technical stakeholder communication, supporting real-time incident management, impact assessment, and operational improvements in a fast-paced, highly available environment. You will work closely with engineering and operational teams to maintain service reliability, improve incident processes, and drive automation initiativesResonsibilities:Monitor, triage, and coordinate responses to production incidents and operational alerts.Act as a central coordination point between engineering teams and key stakeholders during incidents.Assess incident impact, determine severity, and coordinate communications according to SLA commitments.Manage incident lifecycles from detection through resolution and post-incident activities.Maintain external-facing incident communications and status updates.Support incident reporting, root cause analysis (RCA), and operational reviews.Contribute to process improvements, automation initiatives, and operational tooling enhancements.Collaborate with engineering teams to improve observability, monitoring, and incident response capabilities.Participate in reliability-focused development activities and support operational excellence initiatives.We are looking for:5+ years of hands-on experience in Software Engineering, Site Reliability Engineering (SRE), Production Engineering, Incident Operations, or related technical roles. Strong software development background with recent, demonstrable experience building and maintaining production-grade applications and automation in Python. Experience working in on-call environments with SLA/SLO-driven operational responsibilities.Proven ability to operate effectively during high-severity, real-time production incidents. Solid understanding of distributed systems, cloud-native architectures, and large-scale production environments. Experience troubleshooting complex application, infrastructure, and service reliability issues.Advanced proficiency in Python development, including building automation, tooling, integrations, and operational services. Experience with software engineering best practices including testing, code reviews, CI/CD, and version control. Strong understanding of the Software Development Lifecycle (SDLC) and production reliability engineering principles. Familiarity with Kotlin is a plus. Experience with Slack automation and operational workflows is desirable. Incident Response & Reliability.Hands-on experience coordinating, managing, and resolving production incidents. Experience assessing customer impact, driving remediation efforts, and leading technical investigations. Ability to create and execute operational runbooks and automate repetitive operational tasks. Strong understanding of observability, monitoring, alerting, and incident response processes. Experience performing root cause analysis and driving continuous reliability improvements.Monitoring and observability platforms (e.g., Datadog, Chronosphere)Incident management platforms (e.g., PagerDuty, Rootly)APIs and service integrations Production debugging and root cause analysisReliability engineering concepts including SLI/SLOs, error budgets, toil reduction, and automated remediation This role is not perfectly suited for you, but you have a friend who would fit? Recommend your friend and get up to 5000 zł!Referral Program: Talent from your networkDon't hesitate and apply noI