Clera

Robotics Data Infrastructure Engineer

Los Angeles · Posted 2h ago

salary not listedmidpermanentonsite
PythonC++LiDARETLELTCI/CD

Job Description

About the Role

This role sits at the intersection of robotics and data infrastructure at a well-funded, early-stage robotics company. You'll be responsible for building reliable pipelines and storage systems that make large volumes of robot telemetry and sensor data usable for engineering and ML teams. Your work will directly enable faster iteration and safer robotic systems — from raw sensor ingestion all the way through to training-ready datasets and real-time analytics.

You'll join a cross-functional team of robotics engineers, software engineers, and data scientists in a fast-paced, on-site environment in Los Angeles, CA. This is a high-impact, hands-on role with broad scope at a company building at the frontier of physical AI and robotics.

Please note: Visa sponsorship is not available for this role.

What You'll Do

  • Design and build scalable data pipelines to ingest and process robot telemetry and sensor data (camera, LiDAR, IMU, and more).

  • Implement storage solutions and schemas that support analytics, model training, and data replay.

  • Ensure data quality, validation, and lineage across ingestion and transformation stages.

  • Optimize latency and throughput for both real-time and batch processing use cases.

  • Instrument observability, monitoring, and alerting for data flows and infrastructure.

  • Collaborate closely with robotics engineers and data scientists to translate platform needs into production-grade implementations.

  • Productionize ETL/ELT workflows with CI/CD and automated testing.

  • Troubleshoot and resolve production incidents affecting data availability or correctness.

What We're Looking For

Required:

  • 3+ years of hands-on experience building data infrastructure or engineering pipelines specifically for robotics sensor data — this is a dealbreaker requirement.

  • Proven experience designing, building, and maintaining data ingestion, processing, and storage pipelines for sensor data (e.g., camera, LiDAR, IMU).

  • Strong fundamentals in distributed systems, databases, and data pipeline design.

  • Proficiency in Python and/or C++ for building data tooling and pipelines.

  • Hands-on experience with cloud data platforms and distributed processing tools — e.g., AWS or GCP, Kafka or Pub/Sub, Spark or Flink, Airflow.

  • Experience with containerization and deployment of data pipelines using Docker and Kubernetes, plus basic CI/CD.

  • Experience with time-series databases and telemetry data management in a robotics context.

  • Strong communication skills and a collaborative mindset for working across engineering and ML teams.

Nice to Have:

  • Experience with ROS / ROS2 robotics middleware.

  • Familiarity with ML workflow tooling such as MLFlow or Kubeflow for end-to-end robotics data pipelines.

  • Experience with robotics simulation tools (e.g., Gazebo) and synthetic data generation.

  • Prior experience in an early-stage or high-growth startup environment.

Location

This is a full-time, on-site role based in Los Angeles, CA. Candidates based in or willing to relocate to the Los Angeles area are strongly preferred. The company also has a presence in New York City, NY and San Francisco, CA.

Compensation & Benefits

Compensation will be competitive and commensurate with experience, including equity participation appropriate for an early-stage company. Specific details will be shared during the interview process.