Craftware

Site Reliability Engineer

Remote · Posted 16h ago

salary not listedseniorfreelanceremote
SalesforceVeevaUiPathDatabricks

Job Description

Craftware is a technology company of over 500 experts, empowering large organizations to solve complex business challenges with modern IT solutions - from sales systems and automation to data platforms and AI. We operate where technology must be reliable, secure, and scalable. We deliver end-to-end projects: from analysis and architecture through implementation to development and maintenance. We are a trusted partner of industry leaders such as Salesforce, Veeva, UiPath, and Databricks.Model: Remote Engagement: Full-time (B2B)We are looking for an experienced Site Reliability Engineer (SRE) to ensure the reliability, availability, performance, and operational continuity of a complex enterprise ecosystem consisting of multiple highly integrated systems. The SRE will be responsible for the technical stability of individual applications as well as the end-to-end reliability of critical business processes spanning multiple platforms and services, working closely with DevOps teams, application owners, architects, support teams, and Single Points of Contact (SPOCs) responsible for integrated and dependent systems.ResponsibilitiesEnsure high availability, reliability, performance, and operational continuity of business-critical systems and servicesMonitor and analyze end-to-end business processes across multiple applications, APIs, middleware components, messaging platforms, and external systemsIdentify system dependencies and assess their impact on business service availabilityEstablish and maintain monitoring, observability, alerting, and dashboards covering both technical and business-process metricsDefine and track SLIs, SLOs, availability targets, and other reliability metricsCoordinate major incidents involving multiple systems and technical teamsLead root cause analysis for production incidents, integration failures, performance degradation, and service interruptionsCoordinate troubleshooting and communication with SPOCs, system owners, vendors, infrastructure teams, and external providersManage and prioritize the work of the DevOps team responsible for deployment, monitoring, automation, infrastructure, and operational supportDrive automation of operational activities, deployments, health checks, recovery procedures, and system maintenanceMaintain operational runbooks, troubleshooting guides, escalation paths, and recovery proceduresEnsure appropriate backup, disaster recovery, failover, and business continuity mechanisms are implemented and validatedSupport release planning, production readiness, risk assessment, dependency analysis, and rollback strategiesProactively identify reliability risks, performance bottlenecks, single points of failure, and architectural weaknessesWork with development and architecture teams to improve resilience, scalability, fault tolerance, retry mechanisms, and graceful degradationLead post-incident reviews and ensure corrective and preventive actions are implementedRequirementsStrong experience in Site Reliability Engineering, DevOps, Production Engineering, Application Operations, or a similar roleExperience working with complex, highly integrated enterprise architecturesStrong understanding of end-to-end business process monitoring and dependency managementExperience with incident management, root cause analysis, problem management, and service restorationPractical knowledge of monitoring, logging, alerting, and observability platformsGood understanding of SLI, SLO, SLA, availability, latency, throughput, and reliability conceptsExperience with CI/CD, release management, infrastructure automation, and deployment processesUnderstanding of high availability, disaster recovery, failover, and resilience patternsAbility to coordinate technical activities across multiple teams and system ownersExperience managing or coordinating a DevOps or operations-focused engineering teamStrong analytical, troubleshooting, and communication skillsNice to haveExperience with cloud platforms (AWS, Azure, or GCP) in a production operations contextExperience with containerization and orchestration (Docker, Kubernetes)Scripting/automation skills (Python, Bash, or similar)Familiarity with APM/observability tooling (e.g., Datadog, Dynatrace, New Relic, Grafana, Prometheus)Experience with messaging/integration middleware (e.g., Kafka, MQ, ESB platforms)ITIL or similar IT service management framework knowledgeWe offerB2B contract (rate up to 190 PLN net/h + VAT)Fully remote service deliveryBroad range of projects (internal, international) - genuine variety of clients and tasksBudget for skills development and certifications as part of the collaborationRegular collaboration reviews and discussion of project scopeAdditional benefits available as part of the collaborationNetworking and team-building events for project teams