Lyrebird Health
Staff Site Reliability Engineer
Melbourne · Posted 3h ago
Job Description
Staff Site Reliability Engineer
The Role
We're hiring a Staff SRE to own the reliability, security and scalability of Lyrebird's production platform. That covers our AWS infrastructure, deployment pipelines, observability, and how we detect and respond to incidents. You'll report to our Engineering Manager, Platform and Security, and work across every engineering team to make our services resilient and easy to operate. Clinicians use Lyrebird during patient consultations, so when the platform has a problem, they feel it straight away.
This is a hands-on technical leadership role. You'll still write the Terraform and dig into production issues yourself, but you'll also set the standards other engineers build to: how we define SLOs, how we ship code safely, and how teams run their own services without relying on the platform team for everything. It suits someone who has owned infrastructure end to end, stays steady making calls during an incident, and raises it early when something isn't safe.
About Us
Every day in Australia, 10 people die from preventable medical errors. Not from incurable diseases, but from misdiagnosis, missed information, and clinicians making decisions under impossible pressure. At the centre of that failure is a doctor spending 40% of their day on paperwork instead of patients. Lyrebird is fixing that.
We turn real clinical conversations into accurate notes, automatically, backed by ambient real time document creation, intelligent document management and clinical workflow automation, native patient data integrations across siloed health systems, and a patient and clinical communications layer. Thousands of clinicians across multiple disciplines use Lyrebird every day in environments where accuracy, safety, and time matter deeply. We're growing fast across Melbourne, the UK, and Asia, and there's a finite window right now to set the standard for clinical AI in Australia before anyone else does.
What you'll do
Define the SLOs, alerting and observability strategy that show how the platform is really performing, and hold the line on availability and latency for the clinicians who depend on it
Shape how we detect, respond to and learn from incidents, so customer-impacting issues become rarer and get resolved faster when they do happen
Evolve our AWS and Terraform platform for resilience, scale and disaster recovery, staying ahead of our growth rather than catching up to it
Make deployments safer and faster by setting the direction for our CI/CD pipelines and developer tooling
Close security and compliance gaps by building controls and compliance automation into the platform, not bolting them on afterwards
Raise operational maturity across engineering so teams can run their own services confidently, with far less manual support from platform
Challenge technical decisions when reliability or security is at stake, and bring the evidence to back it up
What you'll bring
Deep hands-on experience running production workloads on AWS with Terraform, containers and CI/CD, where getting it wrong had real consequences
A strong grounding in networking, security and production troubleshooting, the kind that finds the root cause when the dashboards disagree
A track record of designing observability, defining SLOs and leading incident response that measurably improved reliability
The coding ability to automate operational work away rather than absorb it
Evidence of influencing technical direction across multiple engineering teams without relying on authority
Nice to have
Experience with ECS/Fargate, PostgreSQL, and distributed or asynchronous workloads
Familiarity with TypeScript/Node.js and OpenTelemetry
Exposure to healthcare technology or frameworks such as ISO 27001 or Cyber Essentials Plus
Experience operating platforms across multiple regions
Clinicians use Lyrebird in the room with their patients, in moments where accuracy and attention matter most. The platform you build is what makes that dependable: when it simply works, a doctor gets to look up from the screen and back at the person in front of them. That's the standard you'll be holding, and it's one worth caring about.
We're building a team that reflects the diversity of the people who benefit from our work. We want Lyrebird to be a place where everyone feels safe, supported, and able to thrive. If you're from an underrepresented background in tech, we strongly encourage you to apply, even if you don't meet every requirement.
More jobs like this
Staff Data Platform Engineer
Kayak · Berlin Office
Staff Information Security Professional (w/m/d)
IONOS · Hinterm Hauptbahnhof 3-5, 76137 Karlsruhe
Software Engineer (Staff+ level)
Anthropic · US incl Remote + Canada + Ireland + UK + Switzerland + Japan