Outlier is hiring worldwide — AI Trainers, Coding Experts & Writing EvaluatorsFreelance • Fully remote • $15–$60/hr • Work when you wantBrowse Outlier roles on MMagic.ai →
Outlier is hiring worldwide — AI Trainers, Coding Experts & Writing EvaluatorsFreelance • Fully remote • $15–$60/hr • Work when you wantBrowse Outlier roles on MMagic.ai →
T

Inference Engineer (Remote — Worldwide)

Together AI$200k – $400kRemote (Global)Posted 1w ago
CUDATritonC++20PyTorchLLM InferenceDistributed SystemsGPU ArchitectureFlashAttentionQuantizationSystems Programming

About the role

As an Inference Engineer at Together AI, you will be at the forefront of the generative AI revolution, architecting the world’s fastest and most efficient engine for open-source models. You will design and implement cutting-edge kernels and distributed systems that push the boundaries of tokens-per-second while maintaining strict latency requirements. Your work will directly empower thousands of developers and researchers by making massive-scale LLM deployment economically viable and technically seamless.

Responsibilities

- Optimize and maintain the core Together Inference Engine to achieve industry-leading throughput and latency. - Develop and integrate custom CUDA and Triton kernels to accelerate attention mechanisms and linear layers. - Implement advanced scheduling algorithms and memory management strategies like PagedAttention to maximize hardware utilization. - Research and deploy state-of-the-art quantization and distillation techniques to improve model efficiency without sacrificing quality. - Collaborate with the research team to provide seamless support for new open-source model architectures (Llama 3, Mixtral, DBRX). - Profile and benchmark inference performance across diverse hardware configurations to identify and eliminate bottlenecks. - Build robust, scalable APIs and microservices that interface directly with GPU clusters at massive scale.

Requirements

- Professional experience building and scaling large-scale distributed inference systems for LLMs. - Deep expertise in GPU architecture (NVIDIA Hopper/Ampere) and writing high-performance CUDA kernels. - Strong proficiency in modern C++ and Python, with a focus on systems programming and memory management. - Hands-on experience with optimization libraries such as FlashAttention, vLLM, TensorRT-LLM, or Triton. - Solid understanding of transformer architectures, quantization techniques (FP8, INT8, AWQ), and speculative decoding. - Ability to work independently in a fully remote, globally distributed environment across multiple time zones. - Proven track record of contributing to open-source systems or publishing research in high-performance computing.

Benefits

Equity, benefits, remote-first.
Disclaimer: MMagic.ai connects talented people with AI companies around the world. While we work hard to feature quality opportunities, we don't independently verify employers, candidates, salaries, or hiring outcomes. We encourage you to research each opportunity and company before applying or making an offer.

Similar roles