Site Reliability Engineer
New
S
STN IncCloud Platform Engineering
Remote (US)Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- PythonKubernetesGoGrafanaPrometheusDatadog
Requirements
- 5+ years in SRE, DevOps, or production engineering roles.
- Strong programming skills in Go or Python.
- Hands-on experience operating Kubernetes-based platforms at scale.
- Deep familiarity with observability tooling such as Prometheus, Grafana, Datadog, or OpenTelemetry.
- Strong incident management experience including major-incident command.
Responsibilities
- Define and operate Service Level Objectives (SLOs) aligned with customer SLAs.
- Build and maintain the observability stack including metrics, logs, traces, and alerting.
- Lead incident response and chair post-incident reviews.
- Drive automation to reduce toil and improve mean-time-to-recover (MTTR).
- Author and maintain operational runbooks alongside the NOC.
- Manage on-call rotation, escalation paths, and incident-management tooling.
- Coordinate cross-functionally with NOC, Platform Engineering, and Network Engineering.
- Drive chaos engineering, game days, and reliability testing programs.
- Produce SLA performance reports in coordination with the SLA Manager.
- Mentor junior engineers and contribute to engineering culture.
View Full Description & ApplyYou'll be redirected to the employer's site