A
Inference Performance Engineer
FeaturedCUDAInferencePythonDistributed Systems
About the role
Make Claude faster and cheaper to serve. You will profile and optimize the inference stack end to end, from kernels to scheduling, for models served to millions of users. Fully remote, globally.
Responsibilities
Profile and optimize serving throughput and latency. Improve batching, caching and scheduling. Partner with research on model-serving co-design.
Requirements
Strong systems programming (C++/CUDA/Rust or Go). Experience with GPU inference or high-throughput distributed services.
Benefits
Fully remote worldwide, equity, learning budget, home office stipend.
Disclaimer: MMagic.ai connects talented people with AI companies around the world. While we work hard to feature quality opportunities, we don't independently verify employers, candidates, salaries, or hiring outcomes. We encourage you to research each opportunity and company before applying or making an offer.