Outlier is hiring worldwide — AI Trainers, Coding Experts & Writing EvaluatorsFreelance • Fully remote • $15–$60/hr • Work when you wantBrowse Outlier roles on MMagic.ai →
Outlier is hiring worldwide — AI Trainers, Coding Experts & Writing EvaluatorsFreelance • Fully remote • $15–$60/hr • Work when you wantBrowse Outlier roles on MMagic.ai →
L

ML Infra Engineer (Remote — Worldwide)

Lambda$170k – $240kRemote (Global)Posted 2w ago
KubernetesGo (Golang)NVIDIA CUDAInfiniBand/RDMATerraformDistributed SystemsBare Metal ProvisioningLinux InternalsPyTorch InfrastructureInfrastructure-as-Code (IaC)

About the role

As an ML Infrastructure Engineer at Lambda, you will design and scale the orchestration layer for the world’s largest public GPU cloud dedicated to deep learning. You will work on the core systems that manage thousands of NVIDIA H100s and H200s, ensuring seamless high-performance computing (HPC) workflows for top AI research labs and startups. Your mission is to bridge the gap between bare-metal performance and cloud flexibility, enabling researchers to train massive foundation models with zero downtime and optimized interconnectivity.

Responsibilities

- Design and implement automated provisioning systems for massive GPU clusters and multi-node training environments. - Develop and maintain Kubernetes operators to manage heterogeneous hardware lifecycles and GPU health monitoring. - Optimize network performance for distributed training jobs, focusing on low-latency interconnects and non-blocking fabrics. - Build robust observability and self-healing systems to detect and remediate hardware failures in Real-time. - Collaborate with the hardware engineering team to integrate the latest NVIDIA architectures and high-performance storage into the cloud platform. - Participate in an on-call rotation to ensure the 99.9% availability of Lambda’s GPU Cloud APIs and infrastructure. - Author technical documentation and RFCs for core infrastructure architectural changes.

Requirements

- 5+ years of experience in Infrastructure Engineering, SRE, or DevOps with a focus on high-scale distributed systems. - Deep technical mastery of Kubernetes, including custom controllers, operators, and networking (CNI). - Proficiency in systems programming with Go, Rust, or Python. - Hands-on experience managing NVIDIA GPU drivers, CUDA environments, and InfiniBand or RoCE networking for RDMA. - Strong understanding of Linux internals, virtualization (KVM/QEMU), and storage protocols like NVMe-over-Fabrics. - Experience with Infrastructure-as-Code (Terraform, Pulumi) and bare-metal provisioning tools. - Ability to work asynchronously across global time zones in a high-growth, remote-first environment.

Benefits

Equity, healthcare, GPU credits.
Disclaimer: MMagic.ai connects talented people with AI companies around the world. While we work hard to feature quality opportunities, we don't independently verify employers, candidates, salaries, or hiring outcomes. We encourage you to research each opportunity and company before applying or making an offer.

Similar roles