Clera
Software Engineer, Voice AI
San Francisco · Posted 2h ago
Job Description
About the Role
Join a small, high-velocity Y Combinator-backed robotics AI startup building the "social brain" that enables humanoid robots to listen, speak, and operate naturally alongside people. As a Software Engineer on the Voice AI team, you'll own cloud-side systems that sit at the heart of real-world robot deployments — where low latency, robust pipelines, and production reliability are non-negotiable. This is a high-ownership, on-site role in San Francisco for engineers who want their work to ship to real robots in the field.
What You'll Do
Design and develop cloud-side systems that help robots intelligently understand and respond to people in real time.
Use production humanoid robots to evaluate and test new features in real-world settings.
Iterate on software deployed to customer robots operating in the field.
Improve composite voice pipelines using state-of-the-art speech and language models.
Add and enhance features including knowledge storage/retrieval, smarter turn-taking logic, and automatic personality tuning.
Reduce interaction latency and make human-robot conversations feel more natural.
Maintain cloud backend code primarily written in Python 3.14 and work with a custom UDP-based streaming protocol.
Deploy and operate services at edge locations across GCP and AWS.
What We're Looking For
Must-haves:
3+ years of professional software engineering experience delivering cloud-based features.
Demonstrated experience with Voice AI development and pipelines.
Experience with latency-sensitive real-time media pipelines (e.g., voice, streaming, games).
Experience building knowledge storage/retrieval components and dialog turn-taking logic for voice AI.
Experience delivering software deployed to customer-facing robots or devices in real-world environments.
Strong cloud development background with AWS or GCP and Kubernetes-based production deployments.
Proficiency in Python for backend services; familiarity with UDP-based streaming protocols.
Experience deploying and operating services at edge locations.
Observability-focused mindset — experience instrumenting systems with metrics and monitoring tools (e.g., Prometheus/Grafana, distributed tracing).
Strong CI/CD practices: linting, testing, packaging, and automation.
Nice to haves:
Experience with containerization and Infrastructure as Code (Terraform, Helm).
Latency optimization and performance tuning for real-time systems.
Deep knowledge of speech processing and voice pipeline architectures.
Robotics hardware integration experience (e.g., ROS) or familiarity with humanoid robots.
Compensation & Benefits
Salary: $150,000 – $225,000 USD annually, depending on experience.
Visa sponsorship is available.
Location
This is an on-site role based in San Francisco, California, USA. Remote work is not available for this position.