500K–850K USD / year
Research Engineer / Performance Engineer, RL Distributed Systems
AIHybrid — San Francisco, CA, New York City, NY, Seattle, WA
Published on 2026-09-29
These details were extracted automatically from the original listing. They may not be complete or fully up to date — it's worth checking the original job post.
About this role
As a Research Engineer on the Distributed Systems team, you'll design and operate systems that support reinforcement learning at scale. The role involves addressing performance bottlenecks and ensuring fault tolerance while collaborating closely with researchers and performance engineers.
About the company
Anthropic builds reliable and interpretable AI systems, including the Claude family of models, for businesses and developers. The company focuses on AI safety research and practical AI assistants for enterprise and consumer use.
The team
Distributed Systems team within RL Engineering
Stack
PythonRustC++GoKubernetes
What you'll do
- Design, build, and operate distributed systems for RL
- Identify and resolve system limitations
- Implement fault tolerance strategies
- Manage resource allocation and autoscaling
- Ensure observability and automation in operations
- Collaborate with performance engineers
What we're looking for
- Strong software engineering skills in Python
- Experience with large-scale distributed systems
- Understanding of distributed systems fundamentals
- Ability to analyze performance metrics
- Debugging complex failures across many hosts
- Strong written communication skills
Nice to have
- Experience with ML training infrastructure
- Knowledge of several tech stack layers
- Building autoscaling solutions
- Experience with container orchestration
- High-performance networking knowledge
- Familiarity with async Python frameworks
Benefits
- Competitive compensation
- Generous vacation and parental leave
- Equity donation matching
