Outlier is hiring worldwide — AI Trainers, Coding Experts & Writing EvaluatorsFreelance • Fully remote • $15–$60/hr • Work when you wantBrowse Outlier roles on MMagic.ai →
Outlier is hiring worldwide — AI Trainers, Coding Experts & Writing EvaluatorsFreelance • Fully remote • $15–$60/hr • Work when you wantBrowse Outlier roles on MMagic.ai →
M

Platform Engineer

Modal Labs$180k – $260kRemote (Global)Posted 1w ago
RustDistributed SystemsKubernetesLinux InternalsContainerizationeBPFgRPCCloud InfrastructurePython SystemsGPU Orchestration

About the role

As a Platform Engineer at Modal, you will build and scale the high-performance orchestration layer that allows developers to run code in the cloud as easily as on their local machines. You will tackle complex problems involving container start times, cold-start optimization, and high-throughput job scheduling for heavy GPU workloads. Your work will directly impact how thousands of AI companies deploy and scale their models, ensuring our serverless runtime remains the fastest and most reliable in the industry.

Responsibilities

- Design and implement core features for the Modal runtime, focusing on low-latency container orchestration and execution. - Optimize the cold-start performance of our serverless environment to sub-second levels for diverse AI workloads. - Develop and maintain the networking stack that handles high-bandwidth data transfers between object storage and compute nodes. - Architect robust scheduling algorithms to efficiently manage thousands of concurrent GPU and CPU tasks. - Build internal tooling and monitoring to ensure 99.9% uptime for our globally distributed infrastructure. - Collaborate with the product team to design APIs that simplify complex infrastructure patterns for end-users. - Participate in an on-call rotation to support and scale our rapidly growing production environment.

Requirements

- 5+ years of experience building mission-critical distributed systems or high-performance infrastructure. - Deep expertise in systems programming, preferably with Rust, C++, or Go. - Strong understanding of Linux internals, including namespaces, cgroups, and containerization primitives. - Experience managing and scaling Kubernetes clusters or building custom container orchestrators. - Proven track record of improving performance and latency in high-throughput network services. - Familiarity with cloud-native infrastructure (AWS/GCP) and infrastructure-as-code principles. - Ability to work autonomously in a fully remote, fast-paced startup environment.
Disclaimer: MMagic.ai connects talented people with AI companies around the world. While we work hard to feature quality opportunities, we don't independently verify employers, candidates, salaries, or hiring outcomes. We encourage you to research each opportunity and company before applying or making an offer.