Site Reliability Engineer (SRE)
New
B
Bright Vision TechnologiesSoftware Development
100% Remote (U.S.)Full-TimeSenior
Salary100,000 - 180,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 10+ years
- Required Skills
- PythonJavaKubernetesGoGrafanaPrometheusCI/CDLinuxDistributed Systems
Requirements
- Bachelor’s degree in Computer Science, Engineering, or related technical discipline.
- 10+ years of SRE, DevOps, or production engineering experience supporting large-scale distributed systems.
- Strong programming skills in Python, Go, or Java.
- Deep, hands-on experience operating Linux at scale.
- Production experience operating Kubernetes and container-based workloads.
- Strong knowledge of observability tooling like Prometheus, Grafana, OpenTelemetry, or ELK/EFK.
- Hands-on experience designing and operating CI/CD pipelines.
- Solid understanding of distributed system design.
- Demonstrated experience leading incident response and conducting post-incident reviews.
- Excellent communication and documentation skills.
Responsibilities
- Define, instrument, and refine service-level objectives (SLOs), SLIs, and error budgets.
- Lead incident response and resolution, facilitating post-incident reviews.
- Design and implement comprehensive monitoring and observability strategies using tools like Prometheus and Datadog.
- Build and maintain robust on-call processes and runbooks.
- Automate operational toil using production-grade scripts in Python, Go, or Bash.
- Architect and operate large-scale Kubernetes clusters and container workloads.
- Design CI/CD pipelines supporting safe and frequent releases.
- Lead capacity planning, performance engineering, and chaos experiments.
- Partner with developers to embed reliability, failure-mode analysis, and dependency hardening.
- Mentor engineers and foster a culture of operational excellence.
View Full Description & ApplyYou'll be redirected to the employer's site