200K–270K USD / year

Staff Site Reliability Engineer

DevOpsRemote — US
Published on 2026-09-25
These details were extracted automatically from the original listing. They may not be complete or fully up to date — it's worth checking the original job post.

About this role

Horizon3 is looking for a Staff Site Reliability Engineer to develop reliability strategies and set standards for their NodeZero platform. This foundational role involves improving operational readiness, incident response, and defining observability practices across engineering teams.

About the company

Horizon3.ai is a technology company specializing in cybersecurity solutions, focusing on automated penetration testing and threat detection for enterprises. They provide a range of services aimed at enhancing security posture and risk management for organizations across various sectors.

Stack

PythonTerraformAWSKubernetesDatadogNew RelicGrafana

What you'll do

  • Develop and evolve SRE strategy and standards
  • Lead cross-functional teams to enhance reliability
  • Establish meaningful SLIs and SLOs
  • Define observability standards
  • Set standards for dashboards and alerts
  • Drive complex reliability initiatives
  • Improve incident management practices
  • Participate in a 24/7 on-call rotation

What we're looking for

  • Experience with large scale distributed systems
  • Deep knowledge of reliability engineering and incident management
  • Experience establishing SLIs and SLOs
  • Backend system development and automation skills
  • Leadership experience during high-severity incidents
  • Excellent communication skills

Benefits

  • Diverse and inclusive culture
  • Career growth opportunities
  • Flexible vacation policy
  • Generous parental leave
  • Equity options
View original job post