M
About the role
Help scale our open-weight and commercial models across regions with low latency and high reliability. Fully remote across time zones.
Responsibilities
Own the serving layer. Improve token throughput. Build autoscaling and observability.
Requirements
Kubernetes, GPU scheduling, and Python or Rust. Experience with vLLM or TensorRT-LLM is a plus.
Benefits
Fully remote, equity, flexible hours, hardware budget.
Disclaimer: MMagic.ai connects talented people with AI companies around the world. While we work hard to feature quality opportunities, we don't independently verify employers, candidates, salaries, or hiring outcomes. We encourage you to research each opportunity and company before applying or making an offer.