Obsidiansecurity
Site Reliability Engineering Lead – Taiwan
Taipei, Taiwan · Posted 2h ago
Job Description
Obsidian Security is the leading SaaS security platform, trusted by global enterprises like Snowflake, T-Mobile, and Algolia. We protect 200+ organizations across North America, Europe, the Middle East, Southeast Asia, Australia, and New Zealand, including many of the world’s largest Fortune 1000 and Global 2000 companies.
Founded in 2017 and backed by top investors like Greylock, Obsidian was built to close a critical gap: securing SaaS apps where business happens—Microsoft 365, Salesforce, and hundreds more. The company does this by offering a complete SaaS security platform to reduce risk, detect and respond to threats, and prevent breaches at the source. Obsidian was built by leaders who redefined endpoint and identity security at CrowdStrike, Okta, Cylance, and Carbon Black. Now, they’re transforming how SaaS is secured.
With AI driving rapid SaaS growth and complexity, agentic AI tools gain privileged access to sensitive data through integrations, creating new risks most security tools miss. Obsidian uniquely detects anomalous OAuth token activity and manages integration risks. Major announcements are on the horizon. Recognizing that SaaS security needs to evolve, Obsidian enables growing organizations to start with a lightweight, prevention-focused browser extension and expand coverage over time.
With global momentum, a growing partner ecosystem including SentinelOne, Databricks, and Google Cloud, and a major fundraise ahead, Obsidian is scaling rapidly toward long-term growth and IPO readiness.
Site Reliability Engineering Lead – Taiwan
About the Role
Obsidian Security is looking for a Site Reliability Engineering Lead to establish and lead production reliability capabilities within our growing Taiwan engineering organization.
You will be responsible for the reliability, security, scalability, and operational effectiveness of cloud services that protect some of the world’s largest enterprises. You will lead work across service reliability, cloud infrastructure, observability, incident management, capacity planning, vulnerability remediation, and production readiness.
This is a hands-on leadership role for someone who can operate effectively during critical incidents while also addressing the engineering and organizational causes behind them. You will build systems and practices that enable product teams to move quickly without compromising availability, security, or customer trust.
You will work closely with the Director of Engineering – Taiwan and global engineering, infrastructure, security, and product teams. As the Taiwan site grows, you will recruit and develop a high-performing SRE team and help create a strong, shared operational culture.
What You’ll Do
-
Establish and lead the SRE function in Taiwan, including its technical roadmap, operating model, hiring plan, and relationship with global teams.
-
Improve the availability, performance, scalability, security, and cost efficiency of Obsidian’s production services.
-
Define service-level indicators, service-level objectives, error budgets, and operational health metrics for critical services.
-
Build and improve observability across applications, data pipelines, APIs, infrastructure, and customer-facing workflows.
-
Lead production incident response, technical coordination, customer-impact assessment, and service recovery.
-
Establish effective on-call practices, escalation paths, operational runbooks, and incident command processes.
-
Facilitate blameless post-incident reviews and ensure corrective actions address systemic causes.
-
Develop automation that reduces manual operations, deployment risk, recovery time, and repetitive work.
-
Partner with engineering teams on production readiness, resilience testing, failure-mode analysis, capacity planning, and safe service rollout.
-
Improve deployment and change-management practices through progressive delivery, automated validation, and reliable rollback mechanisms.
-
Own or coordinate infrastructure vulnerability and CVE management, including exposure assessment, prioritization, remediation, validation, and reporting.
-
Partner with Security and Engineering to strengthen cloud configuration, patching, secrets management, access controls, and infrastructure security.
-
Identify architectural weaknesses and recurring operational issues, then lead cross-team improvements.
-
Measure and reduce operational toil while expanding engineering teams’ ownership of their services.
-
Mentor SREs and software engineers in reliability engineering and production operations.
-
Collaborate with teams across Taiwan, the US, the UK, and Australia to provide effective global operational coverage.
What We’re Looking For
-
Significant experience in site reliability engineering, production engineering, cloud infrastructure, DevOps, or a closely related discipline.
-
Experience leading SRE, infrastructure, or production operations initiatives and mentoring other engineers.
-
Strong hands-on experience operating cloud-native SaaS products or distributed systems in production.
-
Deep knowledge of several of the following:
-
Public cloud infrastructure
-
Containers and orchestration
-
Infrastructure as code
-
CI/CD and deployment automation
-
Monitoring, logging, tracing, and alerting
-
Incident management and production debugging
-
Capacity planning and performance engineering
-
Networking, storage, databases, and distributed systems
-
-
Demonstrated experience defining and using service-level objectives and operational metrics.
-
Experience improving production reliability through engineering and automation.
-
Practical knowledge of infrastructure security, vulnerability management, CVE remediation, and patching processes.
-
Ability to troubleshoot complex failures across services, data systems, networks, and cloud infrastructure.
-
Strong judgment during high-severity incidents and the ability to communicate clearly under pressure.
-
Experience influencing application and platform teams to adopt stronger operational practices.
-
Strong written and verbal communication skills in English.
-
Ability to collaborate effectively across regions, time zones, functions, and cultures.
Nice to Have
-
Experience operating security, identity, data analytics, or other high-volume enterprise SaaS platforms.
-
Experience running large-scale data ingestion and processing infrastructure.
-
Experience establishing a new SRE team or transforming an existing production operations function.
-
Familiarity with compliance and assurance programs relevant to enterprise SaaS.
-
Experience with chaos engineering, resilience testing, automated remediation, or policy-as-code.
-
Experience developing follow-the-sun operational coverage across multiple regions.
-
Experience working in Taiwan or with globally distributed APAC teams.
-
Mandarin proficiency.
What Success Looks Like
-
Taiwan has a strong SRE team that operates as an integrated part of Obsidian’s global reliability organization.
-
Critical services have meaningful service-level objectives, actionable observability, and clear ownership.
-
Production incidents become less frequent, less severe, and faster to resolve.
-
Vulnerabilities and CVEs are assessed and remediated through a measurable, dependable process.
-
Engineering teams can release changes rapidly with controlled operational risk.
-
Recurring operational work is automated, and service teams take increasing ownership of production reliability.
More jobs like this
Outbound GTM Lead / Staffing Industry
HireHawk · Remote — Mexico
Lead Research Specialist (Remote) | Philippines & South Africa
HireHawk · Remote — Philippines · $1-2K/mo
Lead Research Specialist (Remote) | LATAM
HireHawk · Remote — Mexico City, Mexico City, Mexico · $2-2K/mo