Back to jobs

Research Engineer/Scientist - Speech/Audio Machine Learning

Paris

Pay
Salary not listed in the saved posting
Work setup
Unconfirmed
Employment
Full-time — employment source
Employment type Full-time
Read the full posting
Check the employer’s page ↗
Education & alternatives
Qualifications - PhD or MSc in Computer Science with a focus on Deep Learning, Signal Processing, or Computational Linguistics. - Record of Research: Published work in top-tier venues (NeurIPS, ICLR, ICASSP, Interspeech).

About Ifm-Us

We are a dedicated research lab for building, understanding, using, and risk-managing foundation models.

In the employer’s words · Read in context

Job description

View original posting ↗

About the Institute of Foundation Models 
 
We are a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.
 
As part of our team, you’ll have the opportunity to work on the core of cutting-edge foundation model training, alongside world-class researchers, data scientists, and engineers, tackling the most fundamental and impactful challenges in AI development. You will participate in the development of groundbreaking AI solutions that have the potential to reshape entire industries. Strategic and innovative problem-solving skills will be instrumental in establishing MBZUAI as a global hub for high-performance computing in deep learning, driving impactful discoveries that inspire the next generation of AI pioneers.
 
The Role
 
As a Research Scientist specializing in speech/audio machine learning, you will contribute to the design and training of SoTA end-to-end neural speech models. You will be responsible for developing the core intellectual property, moving beyond cascaded ASR → TTS systems toward native audio-to-audio multimodal architectures.

Key responsibilities

  • Architectural Design: Develop novel neural architectures for low-latency speech-to-speech translation and generation (e.g., Diffusion, Flow-matching, Transformer-based audio LLMs). 
  • Loss Function Engineering: Design and implement custom objective functions to optimize prosody (emotions, intelligibility, naturalness). 
  • Experimental Iteration: Conduct large-scale training runs, performing ablation studies on model architecture and tokenization strategies. 
  • Evaluation Frameworks: Establish rigorous internal benchmarks using both objective metrics (WER, MCD) and subjective human-in-the-loop (MOS) testing. 

Qualifications

  • PhD or MSc in Computer Science with a focus on Deep Learning, Signal Processing, or Computational Linguistics. 
  • Record of Research: Published work in top-tier venues (NeurIPS, ICLR, ICASSP, Interspeech). 

Employment type

Full-time

Your next step

Check the employer’s posting for the current role and application details.

Already applied? Track this application

Source & posting history

View original posting ↗

Source notes

Source excerpts

Selected passages from the saved posting. Check the full description for conditions and exceptions.

Pay

No pay amount identified in the saved description.

Location & working pattern

Paris

Working pattern and location restrictions need checking in the full posting.

Work authorization

No clear work-authorization passage found. Eligibility is unconfirmed.

Status in our records
Active
First seen by us
Apr 15, 2026
Recorded sightings
120
Last seen by us
Oct 6, 2026
Employer says posted
Dec 30, 2025

These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.

Report an error
AI answers unavailable

We couldn’t identify enough role detail in this saved description to support an AI answer. Read the full posting