Verified current Job

Senior Voice AI Engineer

ABOUT THE ROLE

Job Remote Full source details
Clera Source published Oct 2, 2026 Verified 4 hours ago
✓ 100% verification score · Source: Clera (ashby) · Always confirm final requirements on the original source.
Complete source information imported The available role or programme description, requirements, benefits and source facts were imported from the public official endpoint and formatted for reading.
EmploymentFull-time
Work modeRemote / location-flexible

Overview

ABOUT THE ROLE

Full job description

ABOUT THE ROLE As a founding engineer on a small conversational AI team, you will own the real-time voice layer, from incoming speech through AI reasoning to spoken responses. You will help make natural, responsive voice interactions work reliably in production, with a focus on end-to-end latency. WHAT YOU'LL DO

  • Build and own streaming speech-to-text, LLM turn-taking, text-to-speech, and telephony or WebRTC transport.
  • Measure and reduce latency, targeting first audio under 800 milliseconds on real calls.
  • Address interruptions, barge-in, silence detection, overlapping speech, poor audio, accents, and mid-sentence changes.
  • Build an evaluation harness from recorded calls, transcripts, and scored turns to detect regressions and guide product decisions.
  • Compare voice providers and models through evidence-based testing, and make changes based on results.
  • Instrument production systems for turn latency, transcription confidence, drop-offs, and cost per minute.
  • Work directly with founders and make technical decisions in a fast-moving team. WHAT WE'RE LOOKING FOR
  • At least 5 years building production software, including 2 or more years shipping voice, speech, or real-time audio systems.
  • Experience building and shipping end-to-end real-time voice pipelines, including streaming speech recognition, LLM turn-taking, speech synthesis, and telephony or WebRTC.
  • Strong Python or TypeScript skills and comfort working in both.
  • Hands-on experience with an audio stack such as LiveKit, Pipecat, Vapi, Twilio Media Streams, Daily, or a custom WebSocket implementation.
  • Experience debugging audio at the frame level, including sample rates, codecs, jitter, and voice activity detection thresholds.
  • Experience building LLM evaluation harnesses, optimizing latency against real-world targets, and using evaluation results to make product decisions.
  • Clear written English for asynchronous communication. Experience with speech model serving or fine-tuning, SIP, telephony, or LLM orchestration frameworks is a plus. COMPENSATION & BENEFITS Compensation is $96,000 USD annually, regardless of location. Visa sponsorship is not available. LOCATION Fully remote, anywhere in the world. Core team overlap is 13:00 to 17:00 UTC.

Tips for this job

Practical JobOpportunity guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.

Verification notes

laptop-ats-crawler v3

Original authoritative source

JobOpportunity is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

Apply through JobOpportunity →

Browse current JobOpportunity listings from Clera (ashby) →

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books