Clera
Staff Engineer — Agentic AI
San Francisco · Posted 1h ago
Job Description
About the Role
This is a senior technical leadership role at the heart of an AI-native engineering software company, owning the core agent intelligence layer that turns mechanical engineers' intent into reliable, cost-efficient multi-step workflows across complex desktop engineering tools. You'll report directly to the CTO and serve as the technical lead for a small team of AI engineers, a user researcher, and domain expert contractors. The work you do here will define real-world product value for enterprise customers.
What You'll Do
Lead development of the agent intelligence layer that executes multi-step workflows across CAD, simulation, and PLM software.
Own the full product loop — from user story definition to implementation to benchmarking against real engineering workflows.
Drive agent task success rate by defining evaluation frameworks, establishing baselines, and iterating on performance.
Set and enforce per-task token budgets and track cost per completed workflow to ensure commercial viability.
Design rigorous, reproducible evaluation infrastructure grounded in validated user stories — think SWE-bench-level rigor applied to engineering workflows.
Lead user story mapping and validation through direct interviews and collaboration with domain experts.
Translate validated user stories into testable evals, closing the loop between user research and benchmarking.
Own agent architecture decisions: tool-calling strategies, state management, error recovery, model routing, and context management.
Act as a player-coach — write production code, review designs, unblock the team, and raise the engineering bar.
Collaborate cross-functionally with integrations, product, and customers during POCs to align agent behavior with real-world usage.
What We're Looking For
7+ years in software engineering, including at least 2 years building and shipping real-world agentic LLM systems (tool calling, multi-step workflows, failure recovery, cost control).
Deep experience with LLM application architecture: model selection, context/window management, retrieval strategies, tool-calling frameworks, and orchestration patterns.
Strong evaluation and benchmarking instincts for agentic systems — task completion rates, cost efficiency, failure mode analysis; familiarity with benchmarks such as SWE-bench, GAIA, or τ-bench is a plus.
Proven track record of shipped AI systems with measurable outcomes — not just demos or prototypes.
Strong Python skills and hands-on familiarity with the LLM tooling ecosystem (function calling, tool use APIs, tracing/observability tools, evaluation frameworks).
Technical leadership experience setting direction for small teams (3–6 engineers) and performing meaningful code review and architecture decisions.
Hands-on background with mechanical engineering software — CAD/CAE/PLM or simulation tooling (e.g. Siemens NX/NXOpen, Teamcenter, CATIA, Creo, SolidWorks, Ansys, Abaqus, or similar) — either as a builder of these tools or as a power user inside an engineering or manufacturing org.
Experience shipping AI/LLM tooling on top of proprietary engineering data or desktop engineering software (e.g. an agent or MCP server over CAD/PLM APIs, RAG over engineering repos or schematics).
Familiarity with enterprise deployment constraints, including behavior on locked-down corporate workstations.
Experience with desktop automation or programmatic control of applications (COM or similar) is a strong plus.
Published work, open-source contributions, or benchmark contributions in agentic AI is a plus.
Compensation & Benefits
Salary range: $160,000 – $250,000 USD annually. Visa sponsorship is not available for this role.
Location
On-site in San Francisco, California, USA.