M
Senior Inference Engineer
CUDAvLLMRustDistributed Systems
About the role
Scale Mistral's inference stack serving millions of requests daily. Optimize vLLM/TensorRT-LLM kernels, quantization, and distributed serving. Remote-first across Europe and worldwide time zones.
Disclaimer: MMagic.ai connects talented people with AI companies around the world. While we work hard to feature quality opportunities, we don't independently verify employers, candidates, salaries, or hiring outcomes. We encourage you to research each opportunity and company before applying or making an offer.