Outlier is hiring worldwide — AI Trainers, Coding Experts & Writing EvaluatorsFreelance • Fully remote • $15–$60/hr • Work when you wantBrowse Outlier roles on MMagic.ai →
Outlier is hiring worldwide — AI Trainers, Coding Experts & Writing EvaluatorsFreelance • Fully remote • $15–$60/hr • Work when you wantBrowse Outlier roles on MMagic.ai →
E

Research Engineer, Audio

Featured
ElevenLabsCompensation undisclosedRemote (Global)Posted 2w ago
PyTorchJAX/FlaxDigital Signal Processing (DSP)Generative AIText-to-Speech (TTS)Distributed TrainingLarge Language Models (LLMs)CUDA OptimizationAudio Synthesis

About the role

As a Research Engineer at ElevenLabs, you will be at the forefront of the generative AI revolution, building the most advanced text-to-speech and audio foundational models in the world. You will bridge the gap between academic paper concepts and production-ready architectures, optimizing large-scale models that can synthesize human-quality speech and sound effects in real-time. Join our globally distributed team to redefine how the world interacts with digital content through emotionally rich, low-latency audio synthesis.

Responsibilities

- Design and implement novel architectures for multi-lingual text-to-speech and expressive voice cloning. - Scale training pipelines for foundational audio models using distributed training techniques and performance optimization. - Develop data curation and preprocessing workflows for massive-scale, diverse audio datasets. - Rapidly prototype and iterate on research ideas from recent literature to improve audio fidelity and latency. - Collaborate with infrastructure engineers to deploy models into a high-throughput production environment. - Optimize inference engines for low-latency delivery across various hardware backends. - Monitor and evaluate model performance using both objective metrics and subjective human-in-the-loop testing.

Requirements

- Master’s or PhD in Computer Science, Machine Learning, or a related quantitative field. - Proficiency in Python and deep learning frameworks, specifically PyTorch and JAX. - Strong mathematical foundation in probability, statistics, and signal processing. - Proven track record of training large-scale generative models (GANs, Diffusion, Transformers). - Experience with high-performance computing clusters and distributed training across hundreds of GPUs. - Deep understanding of modern audio synthesis techniques such as Vocoding, WavNet, or Flow-based models. - Excellent communication skills for a fully remote, asynchronous work environment.
Disclaimer: MMagic.ai connects talented people with AI companies around the world. While we work hard to feature quality opportunities, we don't independently verify employers, candidates, salaries, or hiring outcomes. We encourage you to research each opportunity and company before applying or making an offer.

Similar roles