Verified current Job

Research Scientist - Speech/Audio Machine Learning

About the Institute of Foundation Models 

Job Full source details
Ifm Us Paris Source published Dec 30, 2025 Verified 3 hours ago
✓ 100% verification score · Source: Ifm Us (lever) · Always confirm final requirements on the original source.
Complete source information imported The available role or programme description, requirements, benefits and source facts were imported from the public official endpoint and formatted for reading.
EmploymentFull-time

Overview

About the Institute of Foundation Models 

Full job description

About the Institute of Foundation Models

We are a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.

As part of our team, you’ll have the opportunity to work on the core of cutting-edge foundation model training, alongside world-class researchers, data scientists, and engineers, tackling the most fundamental and impactful challenges in AI development. You will participate in the development of groundbreaking AI solutions that have the potential to reshape entire industries. Strategic and innovative problem-solving skills will be instrumental in establishing MBZUAI as a global hub for high-performance computing in deep learning, driving impactful discoveries that inspire the next generation of AI pioneers.

The Role

As a Research Scientist specializing in speech/audio machine learning, you will contribute to the design and training of SoTA end-to-end neural speech models. You will be responsible for developing the core intellectual property, moving beyond cascaded ASR → TTS systems toward native audio-to-audio multimodal architectures.

Architectural Design : Develop novel neural architectures for low-latency speech-to-speech translation and generation (e.g., Diffusion, Flow-matching, Transformer-based audio LLMs). Loss Function Engineering : Design and implement custom objective functions to optimize prosody (emotions, intelligibility, naturalness). Experimental Iteration : Conduct large-scale training runs, performing ablation studies on model architecture and tokenization strategies. Evaluation Frameworks : Establish rigorous internal benchmarks using both objective metrics (WER, MCD) and subjective human-in-the-loop (MOS) testing.

PhD or MSc in Computer Science with a focus on Deep Learning, Signal Processing, or Computational Linguistics. Record of Research : Published work in top-tier venues (NeurIPS, ICLR, ICASSP, Interspeech).

Tips for this job

Practical Job and Scholarship guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.

Verification notes

laptop-ats-crawler v3

Original authoritative source

Job and Scholarship is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

Ifm Us (lever) ↗

Browse current Job and Scholarship listings from Ifm Us (lever) →

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books