Systems Reliability Engineer
New
B
Bright Vision TechnologiesSoftware Development
100% Remote (United States)Full-TimeSenior
Salary100,000 - 150,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 6+ years
- Required Skills
- PythonJavaKubernetesGoGrafanaPrometheusCI/CDLinuxDistributed Systems
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related technical discipline.
- 5+ years of SRE, DevOps, or production engineering experience supporting large-scale distributed systems.
- Strong programming skills in at least one of Python, Go, or Java.
- Deep, hands-on experience operating Linux at scale, including networking and performance tuning.
- Production experience operating Kubernetes and container-based workloads.
- Working knowledge of observability tooling such as Prometheus, Grafana, or ELK/EFK.
- Hands-on experience designing and operating CI/CD pipelines.
- Solid understanding of distributed system design, consistency models, and failure semantics.
- Proven experience leading incident response and conducting post-incident reviews.
- Excellent communication and documentation skills.
Responsibilities
- Define and refine service-level objectives (SLOs), indicators (SLIs), and error budgets.
- Lead incident response and conduct high-quality post-incident reviews.
- Design monitoring, logging, and tracing strategies using tools like Prometheus, Grafana, or Datadog.
- Automate operational toil by developing tools in Python, Go, or Bash.
- Architect and operate large-scale Kubernetes clusters and container workloads.
- Design CI/CD pipelines to promote safe and frequent releases.
- Lead capacity planning, performance modeling, and chaos engineering activities.
- Partner with development teams to implement reliability and failure-resilience patterns.
View Full Description & ApplyYou'll be redirected to the employer's site