T
About the role
Keep thousands of GPUs healthy, scheduled and fully utilized for training and inference customers. Remote globally.
Responsibilities
Operate large GPU fleets. Automate failure detection and recovery. Tune networking and storage.
Requirements
Linux internals, Kubernetes, Slurm or similar. InfiniBand or NCCL experience a strong plus.
Benefits
Remote worldwide, equity, on-call bonus, hardware budget.
Disclaimer: MMagic.ai connects talented people with AI companies around the world. While we work hard to feature quality opportunities, we don't independently verify employers, candidates, salaries, or hiring outcomes. We encourage you to research each opportunity and company before applying or making an offer.