Surf Ai
Senior Software Engineer - Data Platform Team
Tel Aviv · Posted 4h ago
Job Description
Surf AI is the agentic operations platform for enterprise security teams. We don't just surface risk, we close it. Our platform connects context across identity, cloud, HR, IT, and SaaS systems, and uses specialized AI agents to drive remediation end-to-end, with human oversight at every step.
We're backed by Accel, Cyberstarts, and Boldstart Ventures, and trusted by Fortune 500 enterprises already deploying Surf in production.
Our team is small and senior, with deep roots in identity, security, and enterprise infrastructure. We work at the intersection of agentic AI and applied security - and we take seriously what it means to build systems that act in real enterprise environments.
Who are we looking for?
We're looking for a Senior Software Engineer to join our Data Platform Team.
In this role, you’ll design, build, and scale the data platform that processes and analyzes massive volumes of events and connections every day. Your work will power our cybersecurity research and AI capabilities, supporting batch and real-time data processing, complex data modeling, machine learning pipelines, and applications built on LLMs.
This is a hands-on role with ownership across platform design, distributed systems, and production operations. Your initial focus will be strengthening our shared pipeline infrastructure and expanding tooling across the ML lifecycle.
What you'll do
- Build reusable platform capabilities: Develop shared frameworks and tools for ingestion, transformation, and orchestration, working closely with the Research teams.
- Enable research and ML workflows: Build a platform for researchers independently to prepare data, train and evaluate models, and run batch inference.
- Develop data models and services: Define entities, relationships, and data contracts that make diverse source data easier to use across research, engineering, and product.
- Own database operations: Manage company databases, including backups, recovery, access controls, and isolation between customers.
- Own production reliability: Improve data quality and monitoring, troubleshoot failures, manage backfills, and participate in on-call.
- Raise engineering standards: Lead design and code reviews, improve testing and deployment practices, and keep the platform maintainable.
- Improve performance and cost: Optimize distributed processing, database queries, storage, and compute usage.
What you’ll bring
- Strong software engineering skills, with experience designing, building, testing, and operating production systems.
- Hands-on experience with distributed systems, including troubleshooting failures and improving reliability and performance.
- Strong experience with Python and SQL, and experience working with Spark or similar distributed processing technologies.
- Solid understanding of databases, data modeling, and data platform fundamentals.
- Experience owning systems from design through deployment and ongoing production operation.
- Ability to translate research, data science, and engineering needs into reusable platform capabilities.
- Experience with cloud infrastructure and modern deployment practices, ideally AWS, Kubernetes, infrastructure as code, and automated deployments.
- Willingness to learn new technologies and work across application code, data processing, and infrastructure.
Nice to have
- Experience with ClickHouse, PostgreSQL, Dagster, Airflow, or dbt.
- Experience supporting ML workflows or working closely with researchers and data scientists.
- Familiarity with MLflow, Delta Lake, or Unity Catalog.
- Experience building platforms that serve multiple customers.
- Background in identity, cybersecurity data, or modeling relationships across systems.
Why Join Us?
If you want to work on foundational systems, ship AI into production, and help define how agentic security actually operates, this is an opportunity to do it early and with real ownership.