Overview
ABOUT THE ROLE
Full job description
ABOUT THE ROLE This is an end-to-end ownership role for a cloud-based ASR and transcription pipeline at an early-stage ambient intelligence consumer startup. You'll work directly with product and general management leadership as one of the company's first US engineering hires, making real tradeoffs between latency, accuracy, and reliability as the product evolves. WHAT YOU'LL DO
- Build and iterate on the cloud-based ASR pipeline, from audio capture through post-processing, in production at scale.
- Own ASR quality and reliability end-to-end, shipping measurable improvements on latency, small-word accuracy, and voice-print reliability.
- Work across data, training and fine-tuning, evaluation, and deployment to turn product feedback into shipped pipeline changes.
- Collaborate closely with overseas R&D, hardware, and supply-chain teams across time zones.
- Partner with a product engineer on shared backend and pipeline surfaces.
- Operate with minimal specification, translating lightweight asks into concrete, production-ready improvements. WHAT WE'RE LOOKING FOR
- 3+ years building and tuning transcription and ASR pipelines end-to-end in production, primarily in cloud-based settings.
- Demonstrated ownership of production ASR systems through the full lifecycle: data preparation, model training and fine-tuning, evaluation, and deployment.
- Hands-on experience with latency-sensitive or streaming audio and ASR pipelines.
- Proficiency across the ML lifecycle, including data handling, evaluation metrics, and production deployment.
- Track record of debugging and tuning transcription quality issues such as small-word accuracy, voice-print reliability, and latency.
- Experience in early-stage or founding engineering environments, shipping without large team support or fully-specified requirements.
- On-device or embedded ML experience (Core ML, TensorFlow Lite, or similar frameworks) is a plus.
- Prior experience with wearable, hardware, or robotics products is a plus.
- Background at AI-native consumer applications focused on transcription or audio is a plus.
- Experience building agent or LLM-based product features, including tool use, memory, or retrieval, is a plus.
- You care about how transcription feels to use, not just how it benchmarks, and can make latency and accuracy tradeoffs independently. COMPENSATION & BENEFITS Salary range: $150,000 to $200,000 USD annually. Visa sponsorship is not available for this role. LOCATION Hybrid, 3 days per week in office. San Francisco Bay Area, California, United States.
Tips for this job
Practical Job and Scholarship guidance. These tips do not replace official rules or create new eligibility requirements.
- Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
- Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
- Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
- Apply through the original employer or official recruitment destination shown on this page.
Verification notes
laptop-ats-crawler v1
Job and Scholarship is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.
Clera (ashby) ↗Browse current Job and Scholarship listings from Clera (ashby) →