250K–325K USD / year

Staff Site Reliability Engineer

DevOpsRemote — United States
Published on 2026-09-18
These details were extracted automatically from the original listing. They may not be complete or fully up to date — it's worth checking the original job post.

About this role

Replit is looking for a Staff Site Reliability Engineer to enhance the reliability and scalability of its infrastructure. This role involves leading incident management, implementing observability solutions, and automating operational tasks. The engineer will collaborate with various teams to maintain high reliability standards while scaling Replit's services.

About the company

Replit is an online platform that provides integrated coding environments for developers and learners, enabling users to write, run, and collaborate on code in various programming languages. The company caters to a diverse user base, including educators, students, and professional developers.

The team

Site Reliability Engineering (SRE)

Stack

PythonGoKubernetesDockerGCPTerraformPulumi

What you'll do

  • Architect and implement observability solutions.
  • Define and track service reliability metrics.
  • Lead incident management and response.
  • Drive automation and infrastructure as code improvements.
  • Optimize cloud deployments and performance.
  • Debug and harden distributed systems.
  • Provide guidance on system designs.
  • Mentor and educate the engineering team.

What we're looking for

  • 8-10 years of experience in SRE or similar roles.
  • Strong programming skills in Python or Go.
  • Deep understanding of distributed systems.
  • Experience with Kubernetes and cloud-native technologies.
  • Expertise in monitoring and observability solutions.
  • Strong incident management skills.
  • Knowledge of infrastructure as code tools.
  • Excellent communication and interpersonal skills.

Nice to have

  • Experience with Google Cloud Platform (GCP).
  • Familiarity with observability platforms like Prometheus and Grafana.
  • Background in building high throughput systems.
  • Significant experience with Go and Terraform.
  • Experience in startup environments.
  • Skills in writing training materials.

Benefits

  • Competitive Salary & Equity.
  • 401(k) program with a 4% match.
  • Health, Dental, Vision, and Life Insurance.
  • Flexible Time Off (FTO) + Holidays.
  • Monthly Wellness Stipend.
  • Autonomous Work Environment.
View original job post