Outlier is hiring worldwide — AI Trainers, Coding Experts & Writing EvaluatorsFreelance • Fully remote • $15–$60/hr • Work when you wantBrowse Outlier roles on MMagic.ai →
Outlier is hiring worldwide — AI Trainers, Coding Experts & Writing EvaluatorsFreelance • Fully remote • $15–$60/hr • Work when you wantBrowse Outlier roles on MMagic.ai →
M

LLM Engineer (Remote — Worldwide)

Mistral AIEUR 180k – EUR 380kRemote (Global)Posted 1w ago
PyTorchLLM TrainingDistributed ComputingCUDA TuningMixture of Experts (MoE)DeepSpeedSlurmPythonTransformersModel Quantization

About the role

Join the core engineering team at Mistral AI to push the frontiers of open-weight large language models. You will play a pivotal role in designing, training, and optimizing our next generation of foundation models, from Mistral 7B to the latest MoE architectures like Mixtral. This remote role offers the unique opportunity to influence the global open AI ecosystem by building efficient, high-performance models that power millions of applications worldwide.

Responsibilities

- Design and implement state-of-the-art pre-training and fine-tuning pipelines for next-gen Mistral models. - Optimize model architectures for maximum throughput and memory efficiency during both training and inference. - Conduct rigorous evaluations and benchmarks to ensure model safety, reasoning capabilities, and multilingual performance. - Collaborate with the infrastructure team to scale training jobs across thousands of GPUs using Slurm or Kubernetes. - Research and apply novel techniques in long-context handling, instruction tuning, and preference optimization (RLHF/DPO). - Package and release open-weight artifacts for the global developer community through Hugging Face and other platforms.

Requirements

- Proven track record of training large-scale transformer models (LLMs) in distributed environments. - Deep expertise in PyTorch and hardware-aware optimization techniques like FlashAttention and quantization. - Experience with distributed training frameworks such as Megatron-LM, DeepSpeed, or JAX/XLA. - Solid understanding of Mixture-of-Experts (MoE) architectures and efficient inference serving. - Strong background in high-performance computing (HPC) and managing large GPU clusters (A100/H100). - Contributions to open-source AI projects or a strong portfolio of published research in NLP. - Ability to work independently in a fully remote, fast-paced European startup environment.

Benefits

Equity, health, remote.
Disclaimer: MMagic.ai connects talented people with AI companies around the world. While we work hard to feature quality opportunities, we don't independently verify employers, candidates, salaries, or hiring outcomes. We encourage you to research each opportunity and company before applying or making an offer.

Similar roles