F
Inference Engineer (Remote — Worldwide)
CUDAC++PythonPyTorchLLM InferenceKernel OptimizationNCCLDistributed SystemsTritonGPU Architecture
About the role
Join the team responsible for building the world's fastest inference platform for state-of-the-art generative models. As an Inference Engineer at Fireworks AI, you will optimize large language models (LLMs) and multimodal architectures to achieve industry-leading latency and throughput. You will work at the intersection of systems programming and machine learning to make high-performance AI accessible to developers worldwide.
Responsibilities
- Develop and maintain highly optimized inference kernels for LLMs, Diffusion models, and multimodal architectures.
- Implement cutting-edge research papers to improve prefix caching, speculative decoding, and model distillation.
- Profile and eliminate bottlenecks across the entire inference stack, from Python bindings to low-level GPU memory management.
- Build and scale distributed serving infrastructure to support millions of requests across heterogeneous GPU clusters.
- Collaborate with the product team to integrate new open-source models (Llama, Mixtral, Qwen) within days of their release.
- Optimize memory footprints to enable high-concurrency serving of large-scale models on commodity hardware.
- Participate in an on-call rotation to ensure the reliability and performance of our public API and dedicated deployments.
Requirements
- Deep expertise in CUDA programming and optimizing GPU kernels for NVIDIA architectures.
- Proven track record of optimizing LLM inference using techniques like PagedAttention, continuous batching, and quantization (FP8, INT8, AWQ).
- Experience with high-performance inference frameworks such as vLLM, TensorRT-LLM, or FasterTransformer.
- Strong proficiency in C++ and Python, with a focus on systems-level performance tuning.
- Understanding of distributed systems and collective communication libraries like NCCL for multi-GPU scaling.
- Ability to work independently in a fully remote, globally distributed environment across multiple time zones.
- Master’s or PhD in Computer Science, Accelerated Computing, or a related field, or equivalent industry experience.
Benefits
Equity, healthcare, compute budget.
Disclaimer: MMagic.ai connects talented people with AI companies around the world. While we work hard to feature quality opportunities, we don't independently verify employers, candidates, salaries, or hiring outcomes. We encourage you to research each opportunity and company before applying or making an offer.