500K–850K USD / year

Research Engineer / Performance Engineer, RL Distributed Systems

AIHybrid — San Francisco, CA, New York City, NY, Seattle, WA
Published on 2026-09-29
These details were extracted automatically from the original listing. They may not be complete or fully up to date — it's worth checking the original job post.

About this role

As a Research Engineer on the Distributed Systems team, you'll design and operate systems that support reinforcement learning at scale. The role involves addressing performance bottlenecks and ensuring fault tolerance while collaborating closely with researchers and performance engineers.

About the company

Anthropic builds reliable and interpretable AI systems, including the Claude family of models, for businesses and developers. The company focuses on AI safety research and practical AI assistants for enterprise and consumer use.

The team

Distributed Systems team within RL Engineering

Stack

PythonRustC++GoKubernetes

What you'll do

  • Design, build, and operate distributed systems for RL
  • Identify and resolve system limitations
  • Implement fault tolerance strategies
  • Manage resource allocation and autoscaling
  • Ensure observability and automation in operations
  • Collaborate with performance engineers

What we're looking for

  • Strong software engineering skills in Python
  • Experience with large-scale distributed systems
  • Understanding of distributed systems fundamentals
  • Ability to analyze performance metrics
  • Debugging complex failures across many hosts
  • Strong written communication skills

Nice to have

  • Experience with ML training infrastructure
  • Knowledge of several tech stack layers
  • Building autoscaling solutions
  • Experience with container orchestration
  • High-performance networking knowledge
  • Familiarity with async Python frameworks

Benefits

  • Competitive compensation
  • Generous vacation and parental leave
  • Equity donation matching
View original job post