Systems Reliability Engineer

New
B
Bright Vision TechnologiesSoftware Development
100% Remote (United States)Full-TimeSenior
Salary100,000 - 150,000 USD per year
Apply NowOpens the employer's application page

Job Details

Experience
6+ years
Required Skills
PythonJavaKubernetesGoGrafanaPrometheusCI/CDLinuxDistributed Systems

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or a related technical discipline.
  • 5+ years of SRE, DevOps, or production engineering experience supporting large-scale distributed systems.
  • Strong programming skills in at least one of Python, Go, or Java.
  • Deep, hands-on experience operating Linux at scale, including networking and performance tuning.
  • Production experience operating Kubernetes and container-based workloads.
  • Working knowledge of observability tooling such as Prometheus, Grafana, or ELK/EFK.
  • Hands-on experience designing and operating CI/CD pipelines.
  • Solid understanding of distributed system design, consistency models, and failure semantics.
  • Proven experience leading incident response and conducting post-incident reviews.
  • Excellent communication and documentation skills.

Responsibilities

  • Define and refine service-level objectives (SLOs), indicators (SLIs), and error budgets.
  • Lead incident response and conduct high-quality post-incident reviews.
  • Design monitoring, logging, and tracing strategies using tools like Prometheus, Grafana, or Datadog.
  • Automate operational toil by developing tools in Python, Go, or Bash.
  • Architect and operate large-scale Kubernetes clusters and container workloads.
  • Design CI/CD pipelines to promote safe and frequent releases.
  • Lead capacity planning, performance modeling, and chaos engineering activities.
  • Partner with development teams to implement reliability and failure-resilience patterns.
View Full Description & ApplyYou'll be redirected to the employer's site
100,000 - 150,000 USD per year
Apply Now