O
Data Infrastructure Engineer
SparkRayPythonData Engineering
About the role
Build the data pipelines that feed frontier model training: petabyte-scale ingestion, dedup, filtering and quality signals. Remote worldwide.
Responsibilities
Own large-scale batch and streaming pipelines. Build dataset quality tooling. Optimize storage and cost.
Requirements
Experience with Spark, Ray or Beam. Strong Python. Petabyte-scale data experience.
Benefits
Remote worldwide, equity, learning budget.
Disclaimer: MMagic.ai connects talented people with AI companies around the world. While we work hard to feature quality opportunities, we don't independently verify employers, candidates, salaries, or hiring outcomes. We encourage you to research each opportunity and company before applying or making an offer.