G
Inference Software Engineer (Remote — Worldwide)
FeaturedC++PythonCUDA/TritonLLM InferenceComputer ArchitectureDistributed SystemsCompilersPerformance TuningSRAM OptimizationRuntime Systems
About the role
Join the team responsible for maintaining Groq's industry-leading inference speeds by optimizing the software stack that drives our Language Processing Units (LPUs). You will focus on the intersection of compiler outputs and hardware execution to ensure production LLM workloads run with deterministic, ultra-low latency. Your work will directly impact how the world deploys real-time AI applications by eliminating bottlenecks in the inference pipeline.
Responsibilities
- Design and implement high-performance kernels to squeeze maximum throughput out of Groq’s software-defined hardware.
- Optimize the hand-off between the Groq compiler and the runtime execution engine for multi-chip LPU systems.
- Develop and maintain low-latency drivers and communication libraries for large-scale model deployments.
- Profile production LLM workloads to identify and eliminate microsecond-level bottlenecks in the inference path.
- Collaborate with the compiler and hardware teams to influence the design of future silicon generations based on software performance data.
- Contribute to the stability and scalability of GroqCloud’s inference engine, supporting thousands of concurrent requests.
Requirements
- 5+ years of experience in systems programming with C++ and Python.
- Deep understanding of computer architecture, specifically regarding SRAM, high-bandwidth memory, and PCIe interconnects.
- Proven track record of optimizing high-performance kernels or compilers for AI accelerators (TPUs, GPUs, or custom ASICs).
- Experience with distributed systems and collective communication primitives like NCCL or RCCL.
- Strong knowledge of Transformer architecture internals and KV-cache management techniques.
- Ability to work asynchronously across global time zones while maintaining high code quality and documentation standards.
- Experience with low-level performance profiling tools (e.g., Nsight, VTune, or custom hardware telemetry).
Benefits
Equity, premium health, learning budget.
Disclaimer: MMagic.ai connects talented people with AI companies around the world. While we work hard to feature quality opportunities, we don't independently verify employers, candidates, salaries, or hiring outcomes. We encourage you to research each opportunity and company before applying or making an offer.