Overview
ABOUT THE ROLE
Full job description
ABOUT THE ROLE This is a senior technical leadership position at the heart of an enterprise AI platform for hardware engineers, responsible for the agent intelligence layer that turns real engineering intent into reliable, multi-step automated workflows across desktop CAD, simulation, and PLM tools. You will report directly to the CTO and serve as the technical lead for a small team of AI engineers, a user researcher, and domain expert contractors. The work you do here will define the product's real-world value to enterprise customers. WHAT YOU'LL DO
- Own the core agent intelligence layer that executes multi-step workflows across complex desktop engineering software.
- Drive agent task success rate by defining evaluation frameworks, establishing baselines, and iterating systematically on completion metrics.
- Set and enforce per-task token budgets, tracking cost per completed workflow to ensure commercial viability.
- Design rigorous, reproducible evaluation infrastructure grounded in validated real user stories rather than synthetic tasks.
- Lead user story mapping and validation through direct interviews with engineers and close collaboration with domain experts.
- Translate every validated user story into a concrete test case, closing the loop between user research and agent benchmarking.
- Make foundational architecture decisions covering tool-calling strategies, state management, error recovery, model routing, and context management.
- Act as a player-coach: write production code, review architecture decisions, unblock teammates, and raise overall engineering standards.
- Collaborate cross-functionally with integrations, product, and customers during proofs-of-concept to align agent behavior with real-world usage. WHAT WE'RE LOOKING FOR
- 7+ years of software engineering experience, with at least 2 years building LLM-based agents that take real-world actions.
- Exceptional technical depth in agentic systems, including model selection, context and window management, retrieval, tool calling, and orchestration patterns.
- Demonstrated experience building evaluation and benchmarking frameworks that measure task completion, cost efficiency, and failure modes.
- Strong Python proficiency and hands-on familiarity with LLM tooling: function calling, tool APIs, observability and tracing, and evaluation frameworks.
- Proven track record shipping AI or LLM tooling on top of proprietary engineering data or desktop engineering software, such as agents or MCP servers over CAD, PLM, simulation, or similar systems. General-purpose chatbot or web-app RAG work alone is not sufficient.
- Experience setting technical direction and reviewing code for small engineering teams while continuing to write production code yourself.
- Experience with desktop automation or programmatic application control, such as COM or similar interfaces.
- Background in mechanical engineering, CAD/CAE, PLM, or an adjacent engineering-software domain.
- Familiarity with deploying agents on locked-down enterprise workstations and the associated security and operational constraints.
- Published work, public benchmarks, or open-source contributions in agentic AI are a strong plus. COMPENSATION AND BENEFITS Base salary: $160,000 to $250,000 USD annually, plus equity. Visa sponsorship is not available for this role. LOCATION On-site in San Francisco, California, United States.
Tips for this job
Practical JobOpportunity guidance. These tips do not replace official rules or create new eligibility requirements.
- Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
- Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
- Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
- Apply through the original employer or official recruitment destination shown on this page.
Verification notes
laptop-ats-crawler v3
JobOpportunity is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.
Clera (ashby) ↗