Outlier is hiring worldwide — AI Trainers, Coding Experts & Writing EvaluatorsFreelance • Fully remote • $15–$60/hr • Work when you wantBrowse Outlier roles on MMagic.ai →
Outlier is hiring worldwide — AI Trainers, Coding Experts & Writing EvaluatorsFreelance • Fully remote • $15–$60/hr • Work when you wantBrowse Outlier roles on MMagic.ai →
C

ML Engineer — Code Generation

Codeium$180k – $280kRemote (Global)Posted 2w ago
PyTorchLLMsDistributed TrainingCUDAC++PythonNLPTransformersModel QuantizationDeepSpeed

About the role

As an ML Engineer on the Code Generation team, you will be at the forefront of building the most advanced AI coding assistants, including our flagship Windsurf IDE and the Codeium extension suite. You will focus on scaling and refining large language models (LLMs) to handle complex, multi-file reasoning and real-time completions. Your work will directly impact millions of developers by pushing the boundaries of what is possible in automated software engineering and context-aware code synthesis.

Responsibilities

- Train and fine-tune state-of-the-art transformer models for code completion, generation, and multi-file refactoring. - Develop and implement innovative context-window management strategies to allow models to reason over massive, private codebases. - Optimize model latency and throughput to ensure a seamless 'Flow' experience within the Windsurf IDE. - Design and execute rigorous evaluation pipelines to measure model performance across diverse programming languages and architectural patterns. - Collaborate with the systems team to integrate models into our proprietary inference engine for maximum hardware utilization. - Research and adopt emerging techniques in RLHF and DPO to align model outputs with developer intent and best practices. - Improve the data flywheel by curating high-quality synthetic and open-source datasets for specialized code tasks.

Requirements

- 3+ years of professional experience training and fine-tuning Large Language Models (LLMs) for production environments. - Deep expertise in PyTorch and experience with distributed training frameworks like DeepSpeed, Megatron-LM, or FSDP. - Proven track record of working with code-specific datasets and understanding the nuances of abstract syntax trees (ASTs) and repository-level context. - Strong proficiency in Python and C++ for high-performance model serving and optimization. - Experience with inference optimization techniques such as quantization (FP8/INT8), KV caching, and speculative decoding. - Ability to work independently in a fully remote, fast-paced environment with a bias toward shipping production-ready code. - Advanced degree (Masters or PhD) in Computer Science, Mathematics, or a related field with a focus on Deep Learning.
Disclaimer: MMagic.ai connects talented people with AI companies around the world. While we work hard to feature quality opportunities, we don't independently verify employers, candidates, salaries, or hiring outcomes. We encourage you to research each opportunity and company before applying or making an offer.

Similar roles