T
Inference Infrastructure Engineer
GoRustKubernetesCUDADistributed SystemsPyTorchLLM InferenceLoad BalancinggRPCNVIDIA Triton
About the role
Join the team building the world's fastest inference engine for large language models. As an Inference Infrastructure Engineer at Together AI, you will scale our distributed systems to support thousands of concurrent customers running open-source models like Llama 3, Mixtral, and Qwen. You will play a pivotal role in optimizing our custom inference stack, ensuring low-latency delivery and high throughput across thousands of GPUs.
Responsibilities
- Architect and maintain the distributed system that powers the Together Inference API and Custom Models.
- Optimize GPU resource allocation and autoscaling logic to improve cluster utilization and reduce cold starts.
- Integrate and benchmark leading-edge open-source models, ensuring optimal performance on our flash-attention kernels.
- Develop robust monitoring and self-healing mechanisms for a global fleet of H100 and A100 nodes.
- Collaborate with the research team to productionize breakthroughs in quantization and speculative decoding.
- Troubleshoot complex distributed systems failures across the entire stack, from IB networking to Python-based API layers.
Requirements
- 5+ years of experience building distributed systems and high-throughput backend infrastructure.
- Deep expertise in high-concurrency programming using Go, Rust, or C++.
- Proven experience managing large-scale GPU clusters and orchestration (Kubernetes, Slurm).
- Familiarity with deep learning frameworks and inference backends (PyTorch, vLLM, TensorRT-LLM).
- Strong understanding of networking protocols, load balancing, and low-latency API design.
- Experience with cloud-native observability stacks like Prometheus, Grafana, and distributed tracing.
- Ability to thrive in a fast-paced, remote-first startup environment.
Disclaimer: MMagic.ai connects talented people with AI companies around the world. While we work hard to feature quality opportunities, we don't independently verify employers, candidates, salaries, or hiring outcomes. We encourage you to research each opportunity and company before applying or making an offer.