Outlier is hiring worldwide — AI Trainers, Coding Experts & Writing EvaluatorsFreelance • Fully remote • $15–$60/hr • Work when you wantBrowse Outlier roles on MMagic.ai →
Outlier is hiring worldwide — AI Trainers, Coding Experts & Writing EvaluatorsFreelance • Fully remote • $15–$60/hr • Work when you wantBrowse Outlier roles on MMagic.ai →
C

ML Platform Engineer (Remote — Worldwide)

Crusoe$170k – $250kRemote (Global)Posted 1w ago
KubernetesCUDASlurmInfiniBand/RoCEPyTorchGo (Golang)Distributed SystemsHPC (High Performance Computing)NVIDIA Collective Communications Library (NCCL)Terraform

About the role

Join Crusoe’s Infrastructure team to architect the high-performance computing (HPC) software layer that powers large-scale generative AI training on climate-aligned infrastructure. We are building a purpose-built cloud for the LLM era, leveraging stranded energy to fuel massive GPU clusters. As an ML Platform Engineer, you will bridge the gap between bare-metal acceleration and researcher productivity, ensuring seamless orchestration of multi-thousand GPU jobs.

Responsibilities

- Architect and maintain the control plane for massive-scale NVIDIA H100 and B200 clusters across carbon-neutral data centers. - Optimize the job scheduling layer to support massive model training runs involving thousands of interconnected GPUs. - Develop custom Kubernetes operators and telemetry pipelines to monitor GPU health, thermal throttling, and interconnect performance. - Collaborate with AI research customers to debug complex distributed training failures and NCCL bottlenecks. - Build and maintain automated provisioning workflows for high-performance storage solutions like Lustre or Weka. - Implement robust fault-tolerance and checkpointing mechanisms to minimize downtime during multi-week training jobs.

Requirements

- 5+ years of experience in infrastructure engineering, with a focus on ML platforms or distributed systems. - Proven expertise in managing large-scale Kubernetes clusters (EKS, GKE, or bare-metal) and Slurm orchestration. - Deep understanding of high-performance networking fabrics, specifically RoCE v2 or InfiniBand, for low-latency collective communications. - Professional experience with GPU virtualization and container toolkits (NVIDIA Container Toolkit, Enroot, Pyxis). - Proficiency in Go or Python for building platform automation and custom operators. - Familiarity with deep learning frameworks like PyTorch and distributed training libraries such as DeepSpeed or Megatron-LM. - Experience operating in remote-first environments with high degrees of autonomy.

Benefits

Equity, healthcare, remote stipend.
Disclaimer: MMagic.ai connects talented people with AI companies around the world. While we work hard to feature quality opportunities, we don't independently verify employers, candidates, salaries, or hiring outcomes. We encourage you to research each opportunity and company before applying or making an offer.

Similar roles