A
Alignment Research Engineer (Remote — Worldwide)
FeaturedPyTorchRLHFConstitutional AIDistributed TrainingRed TeamingJAXTransformer ArchitecturesModel EvaluationPythonDeep Learning Research
About the role
As an Alignment Research Engineer, you will sit at the intersection of safety theory and large-scale engineering to ensure Claude remains helpful, honest, and harmless. You will develop novel techniques to mitigate risks such as jailbreaking, sycophancy, and power-seeking behavior while improving the model's ability to follow complex instructions. This role is critical to Anthropic’s mission of building frontier models that are robustly aligned with human values even as they scale in capability.
Responsibilities
- Design and implement scalable algorithms for Reinforcement Learning from AI Feedback (RLAIF) to improve model steerability.
- Develop automated safety red-teaming pipelines to identify and patch adversarial vulnerabilities in Claude.
- Conduct empirical research on reward modeling and preference fine-tuning to reduce model bias and hallucinations.
- Build internal tooling and infrastructure to evaluate model alignment across diverse cultural and linguistic contexts.
- Collaborate with the Interpretability team to bridge the gap between model internal states and behavioral safety.
- Monitor and analyze distribution shifts in model behavior during large-scale fine-tuning runs.
- Contribute to whitepapers and technical documentation regarding Anthropic’s safety standards and methodology.
Requirements
- 3+ years of experience in deep learning research or high-performance machine learning engineering.
- Proficiency in Python and experience with large-scale training frameworks like PyTorch, JAX, or TensorFlow.
- Solid understanding of Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI principles.
- Experience working with high-performance computing clusters and distributed training environments.
- Demonstrated ability to implement and iterate on complex research papers or internal safety benchmarks.
- Strong mathematical foundations in probability, statistics, and optimization.
- Proven track record of shipping or publishing work in ML safety, interpretability, or model evaluation.
Benefits
Equity, health, generous PTO.
Disclaimer: MMagic.ai connects talented people with AI companies around the world. While we work hard to feature quality opportunities, we don't independently verify employers, candidates, salaries, or hiring outcomes. We encourage you to research each opportunity and company before applying or making an offer.