Outlier is hiring worldwide — AI Trainers, Coding Experts & Writing EvaluatorsFreelance • Fully remote • $15–$60/hr • Work when you wantBrowse Outlier roles on MMagic.ai →
Outlier is hiring worldwide — AI Trainers, Coding Experts & Writing EvaluatorsFreelance • Fully remote • $15–$60/hr • Work when you wantBrowse Outlier roles on MMagic.ai →
R

Model Performance Engineer (Remote — Worldwide)

Replicate$180k – $360kRemote (Global)Posted 1w ago
CUDATritonPyTorchTensorRTvLLMPythonGPU OptimizationDockerPerformance ProfilingQuantization

About the role

As a Model Performance Engineer at Replicate, you will be the primary driver behind making open-source models run faster and more efficiently than anywhere else. You will work at the intersection of deep learning and systems engineering, optimizing the performance of Cog containers and our underlying inference infrastructure. Your work will directly impact thousands of developers by reducing latency and cost for cutting-edge models like SDXL, Llama 3, and specialized LoRAs.

Responsibilities

- Optimize cold-boot times and inference latency for a wide variety of generative AI models hosted on Replicate. - Develop and integrate customized kernels to accelerate specific model architectures using Triton or CUDA. - Benchmark and profile diverse hardware configurations to determine the most cost-effective path for model deployment. - Collaborate with the infrastructure team to improve how Cog packages and executes models within Docker environments. - Provide technical guidance to the community on best practices for model optimization and efficient weights loading. - Research and implement state-of-the-art quantization techniques (AWQ, GPTQ, FP8) to minimize memory footprints.

Requirements

- Deep understanding of GPU architectures (NVIDIA Ampere/Hopper) and CUDA programming. - Proven experience with performance profiling tools like Nsight Systems, PyTorch Profiler, or Triton. - Familiarity with deep learning optimization frameworks such as TensorRT, vLLM, DeepSpeed, or TVM. - Proficiency in Python and C++, with a strong grasp of the PyTorch internal API. - Experience tuning low-level kernels or managing memory bottlenecks for Large Language Models (LLMs). - Strong communication skills for working asynchronously across a global, fully remote team. - A track record of contributing to or maintaining high-performance open-source projects.

Benefits

Equity, benefits, remote.
Disclaimer: MMagic.ai connects talented people with AI companies around the world. While we work hard to feature quality opportunities, we don't independently verify employers, candidates, salaries, or hiring outcomes. We encourage you to research each opportunity and company before applying or making an offer.

Similar roles