B
ML Infrastructure Engineer
KubernetesRustGoCUDADockerDistributed Systems微服务NVIDIA TritonPythonTerraformgRPC
About the role
As an ML Infrastructure Engineer at Baseten, you will be responsible for building the backbone that allows companies to transition from local model exploration to production-grade deployment. You will work on the core orchestration layer that manages containerized inference workloads, cold starts, and autoscaling across heterogeneous GPU clusters. Your work will directly impact how thousands of developers serve state-of-the-art LLMs and diffusion models with low latency and high reliability.
Responsibilities
- Design and implement the orchestration logic for high-performance model serving on Kubernetes.
- Optimize the model cold-start pipeline to ensure near-instant availability of specialized hardware resources.
- Build and scale Baseten’s internal networking layer to handle high-bandwidth inference traffic across multiple providers.
- Develop custom auto-scaling algorithms that balance cost-efficiency with immediate request availability.
- Collaborate with the product team to design APIs that simplify complex infrastructure tasks for ML practitioners.
- Participate in an on-call rotation to ensure the 24/7 reliability of our global inference fleet.
- Contribute to the open-source ecosystem around model deployment and container orchestration.
Requirements
- 5+ years of experience in infrastructure or systems engineering with a focus on high-traffic production environments.
- Deep expertise in Kubernetes, including writing custom controllers and managing complex networking configurations.
- Proven experience with NVIDIA GPU virtualization, CUDA, and managing hardware-accelerated workloads.
- Strong proficiency in systems programming languages like Rust, Go, or C++.
- Experience building and maintaining large-scale distributed systems, specifically focusing on low-latency data planes.
- Background in optimizing cold start times for containerized applications and managing large Docker image distributions.
- Familiarity with cloud-native storage solutions and high-performance caching strategies.
Disclaimer: MMagic.ai connects talented people with AI companies around the world. While we work hard to feature quality opportunities, we don't independently verify employers, candidates, salaries, or hiring outcomes. We encourage you to research each opportunity and company before applying or making an offer.