Blue
#700 - Incident Operations & Communications Engineer
LatAm 2 · Posted 1h ago
Job Description
BlueCloud is a Snowflake Elite Partner and the 2026 CoCo Catalyst Snowflake Partner of the Year. We help enterprise organizations move from fragmented legacy systems to unified, AI-ready Snowflake platforms — delivering data migration, engineering, governance, BI & analytics, and AI/ML solutions 40–50% faster than traditional approaches. With 450+ Snowflake consultants, 200+ enterprise transformations under our belt, and a 100% Snowflake focus, we combine advisory-led thinking with AI-powered accelerators to turn months of work into weeks of results. Our clients span Financial Services, Healthcare & Life Sciences, Retail, Manufacturing, Energy, and more — and the outcomes speak for themselves: 97% faster reports, 40% fraud reduction, $1.5M in client savings, and 10× client growth. We don't just strategize — we execute. Position Summary: We're seeking a skilled and detail-oriented Incident Operations & Communications Engineer to manage real-time incident response, external communications, and coordination across our production systems. This role sits at the intersection of Software Engineering, Site Reliability Engineering (SRE), Incident Response, and Partner Operations, and is ideal for someone who thrives in a fast-paced, high-stakes environment requiring rapid decision-making under ambiguity. You will play a critical role in the incident lifecycle, including impact assessment, SLA-driven communications, and cross functional coordination with Engineering and Technical Account Managers (TAMs), ensuring timely, accurate, and consistent communication with merchants and partners. This position is a key component to support global operations and enable scalable External Incident Comms as we expand internationally. Key Responsibilities: ● Incident Response & Triage ○ Respond to real-time alerts and determine whether they represent valid incidents. ○ Initiate and participate in incident response workflows. ○ Collaborate with Incident Commanders and different Stakeholders. ○ Manage multiple concurrent incidents while maintaining clarity and prioritization. ● Incident Lifecycle Coordination ○ Act as a central coordination point between: ■ Engineering teams ■ Incident Commanders ■ TAMs and partner-facing teams ○ Support handoffs across time zones as part of the FTS model. ○ Contribute to incident closure, RCA inputs, and reporting workflows, including ownership of SLA merchants and distribution of external SLA Reports. ● Impact Assessment ○ Identify affected merchants and partners using dashboards, alerts, and system signals. ○ Make rapid decisions with incomplete or evolving data. ○ Evaluate incident severity and determine communication requirements based on SLA commitments. ○ Continuously update impact scope as incidents evolve. ● External Communications ○ Own end-to-end communication lifecycle with merchants and partners: ■ Initial notifications within strict SLA windows ■ Ongoing updates aligned with severity-based cadence ■ Final resolution communications ○ Tailor messaging based on merchant-specific requirements and communication rules. ○ Coordinate approvals with stakeholders (TAMs, Engineering, Leadership) when required. ○ Ensure communication quality, clarity, and consistency across all outputs. ● Status Page Management ○ Create and maintain incident entries on merchant-facing status pages. ○ Align internal incident state with external communication. ○ Maintain update cadence based on incident severity (e.g., SEV0, SEV1). ○ Ensure compliance with publishing guidelines and approval processes. ● Operational Scaling & Process Improvements ○ Identify opportunities to reduce manual toil across the incident lifecycle. ○ Contribute to development of: ■ Standardized communication workflows ■ Automation tools and internal systems ■ Improved impact assessment methodologies ○ Collaborate with Engineering and TAMs (NA & EU) to improve: ■ Observability and monitoring ■ Communication decision frameworks ○ Support rollout of new roles and capabilities (e.g., Scribe) within the incident lifecycle. ○ When not on active rotation, contribute directly to core reliability-loop tooling — Scribe, Project Nemo (Deepdive), and Project Elcano — (Python/Kotlin), with hands-on feature ownership, not just identifying improvement opportunities. Qualifications: ● Experience: ○ 4/5+ years of experience in Software Engineering, Incident Operations, DevOps/SRE, or similar roles. ○ Hands-on experience with Python (Kotlin and frontend exposure as a nice to have) ○ Experience working in on-call environments with SLA-driven responsibilities. ○ Familiarity with observability tooling and dashboards and DevOps/SRE concepts ○ Proven ability to operate in high-pressure, real-time incident scenarios ● Technical Background: ○ Strong understanding of distributed systems and production environments. ○ Experience with: ■ Monitoring and alerting systems (e.g., Datadog, Chronosphere, etc.) ■ Incident management tools (e.g., PagerDuty, Rootly, Slack workflows) ○ Familiarity with: ■ APIs and system integrations ■ Understanding of software development lifecycle (SDLC) and production reliability. ● Soft Skills: ○ Strong operational judgment and ability to make decisions under uncertainty. ○ Excellent written and verbal communication skills, especially in external-facing contexts. ○ Ability to manage multiple priorities simultaneously in a high-pressure environment. ○ Highly organized and detail-oriented, with a strong ownership mindset. ○ Strong cross-functional collaboration skills, especially with Engineering and partner teams. ○ Proactive mindset with focus on continuous improvement and scalability