Outlier is hiring worldwide — AI Trainers, Coding Experts & Writing EvaluatorsFreelance • Fully remote • $15–$60/hr • Work when you wantBrowse Outlier roles on MMagic.ai →
Outlier is hiring worldwide — AI Trainers, Coding Experts & Writing EvaluatorsFreelance • Fully remote • $15–$60/hr • Work when you wantBrowse Outlier roles on MMagic.ai →
M

Research Scientist — Long Context

Featured
Magic.dev$280k – $450kRemote (Global)Posted 1w ago
PyTorchCUDAState Space Models (SSM)Linear TransformerDistributed TrainingTritonLarge Language Models (LLMs)Sequence ModelingDeep Learning Theory

About the role

As a Research Scientist focused on Long Context, you will work at the frontier of sequence modeling to build a truly intelligent AI software engineer. Magic is moving beyond the constraints of traditional Transformers to develop novel architectures capable of reasoning over millions of tokens of codebase state. Your work will directly impact our ability to synthesize complex, multi-file software changes and maintain deep coherence across massive context windows.

Responsibilities

- Design and implement novel neural architectures that surpass the scaling limitations of standard self-attention. - Develop advanced evaluation frameworks for measuring long-range reasoning, retrieval accuracy, and code coherence. - Optimize training recipes for models with multi-million token context windows, ensuring efficient memory utilization and throughput. - Collaborate with the systems team to co-design hardware-efficient kernels for proprietary sequence models. - Stay abreast of and synthesize the latest research in efficient ML, applying relevant breakthroughs to our core stack. - Contribute to the fundamental understanding of how long-context models retrieve, weight, and process information from distant tokens.

Requirements

- Ph.D. in Computer Science, Mathematics, or a related field, or equivalent elite-level industry research experience. - Deep expertise in sequence modeling architectures (e.g., State Space Models, Linear Attention, RNNs, or Sparse Transformers). - Proven track record of publishing at top-tier venues (NeurIPS, ICML, ICLR) or building high-impact open-source ML systems. - Strong proficiency in PyTorch and C++/CUDA for implementing custom kernels and performance-critical layers. - Experience managing the stability and hyperparameter tuning of large-scale distributed training runs. - Familiarity with the unique challenges of code-specific data, including abstract syntax trees and repository-level structures. - Ability to work autonomously in a fast-paced, remote-first environment with a high degree of technical ownership.
Disclaimer: MMagic.ai connects talented people with AI companies around the world. While we work hard to feature quality opportunities, we don't independently verify employers, candidates, salaries, or hiring outcomes. We encourage you to research each opportunity and company before applying or making an offer.

Similar roles