Menlo Security

Principal Platform Infrastructure Engineer (SRE Enablement)

EMEA - Distributed (UK) · Posted May 22

salary not listedprincipalpermanentremote
KubernetesGoogle Kubernetes EngineAWSGCPTerraformSpaceliftHelm

Job Description

Menlo Security is the leader in Browser Security for human and agentic workforces. Our mission is to enable humans and agents to connect, communicate, and collaborate securely, without compromise. The Menlo Browser Security Platform protects organizations from cyberattacks by stopping threats across the web, documents, and email before they reach the user. With Menlo Agent Runtime Security (MARS), that protection now extends to the AI agents working alongside every employee. Menlo Security is trusted by major global businesses, including Fortune 500 companies and government agencies, to protect their most valuable asset, their data, and is backed by top-tier investors.

About the Role

Platform Infrastructure Engineering is responsible for building and operating Menlo Security's Infrastructure Platform. Together with the rest of our engineering teams, we enable our customers to connect to the Internet without compromise. Our environment provides services globally. We expect failure, build security in by design, create evolvable systems, and enable multi-tenancy across the infrastructure. Automation is an absolute for us.

We are committed to getting it done properly, the first time.

As a Platform Infrastructure Engineer, you'll join a group of experienced engineers who are part of a globally distributed team responsible for building and managing the company's core infrastructure services and maintaining our constantly growing platform. The team operates a sophisticated cloud-native infrastructure built on Google Kubernetes Engine and VMs spanning multiple environments globally from development to production. We manage infrastructure as code with Terraform and Spacelift orchestration, and deploy services using Helm charts. Our platform emphasizes security-first design, comprehensive observability, and multi-region resilience. Success in this role requires working with a vast VM fleet in AWS and GCP as well as Kubernetes, writing Infrastructure as Code, and a passion for automation and reliability engineering.

Responsibilities

  • Architect and govern the design, deployment, and operation of high-scale, multi-region VM and Kubernetes infrastructure on GCP and AWS, ensuring maximum resilience and performance across all environments.

  • Drive cross-functional technical alignment with Engineering, Product, Compliance, and Security teams, serving as the architectural consultant and leader for major initiatives involving capacity planning, disaster recovery, and cloud-native application design.

  • Define and enforce organizational best practices and standards for Infrastructure as Code (IaC) using Terraform and Spacelift, ensuring consistency and security across all provisioned cloud resources (GCP/AWS).

  • Design and manage complex, multi-layer configuration management and deployment workflows that optimize reliability and operational efficiency across the entire platform.

  • Set the technical direction and implement comprehensive observability solutions (Grafana Cloud, Prometheus/Mimir, OTel collectors), establishing organization-wide standards for system visibility, metrics, and alerting.

  • Define the strategic architecture and lifecycle management of core platform services, including certificate management, DNS automation, ingress controllers, and service mesh networking (Cilium).

  • Proactively identify and lead large-scale strategic efforts to eliminate technical toil and improve operational efficiency through the development of tools, strategic automation, and building advanced CI/CD pipelines.

  • Mentor and provide deep technical guidance to both junior and senior engineers within Platform Infrastructure Engineering.

  • Participate in a 24x7 on-call rotation as part of a globally distributed team, responding to incidents and driving post-incident reviews to ensure long-term solutions and process i

Requirements

  • Bachelor's degree in Computer Science, similar technical field of study, or equivalent practical experience.

  • Proficiency in common programming & scripting languages. We use a lot of python, bash and go.

  • Understanding of network topologies, communication protocols (ie. TCP/IP, HTTP/S, UDP, TLS) and enterprise grade connectivity solutions.

  • Kubernetes expertise including cluster administration, RBAC, networking, workload management, and troubleshooting across production environments.

  • Proven experience with Terraform for infrastructure provisioning and management.

  • Knowledge of Google Cloud Platform services including GKE, VPC networking, Cloud DNS, Artifact Registry, Secret Manager, IAM, Gemini Code Assist, and Workload Identity.

  • Prior experience and success mentoring other junior and senior engineers

  • Experience with GitOps methodologies and tools.

  • Clear understanding of how to use LLM code assist tools to effectively build software.

Follow us on LinkedIn!

Why Menlo?

At Menlo, we don't settle for the status quo — in our technology or our culture. How we think and act is just as important as what we build. Our culture is defined by five core mindsets: Proactive Leadership, Straight Talk, United Impact, Elevated Talent, and Customer-Compelled. We take ownership and drive outcomes without waiting to be told. We communicate directly and seek hard truths. We break down silos and win together. We hold a high bar for ourselves and the people around us. And we treat every customer interaction as mission-critical. If you're someone who sees it, owns it, solves it, and does it — you'll thrive here.

All qualified applicants will receive consideration for employment without regard to race, sex, color, religion, sexual orientation, gender identity, national origin, protected veteran status, or on the basis of disability.

TO ALL AGENCIES: Please, no phone calls or emails to any employee of Menlo Security outside of the Talent organization. Menlo Security’s policy is to only accept resumes from agencies via Ashby (ATS). Agencies must have a valid services agreement executed and must have been assigned by the Talent team to a specific requisition. Any resume submitted outside of this process will be deemed the sole property of Menlo Security. In the event a candidate submitted outside of this policy is hired, no fee or payment will be paid.