Platform Reliability Engineer

New
B
Bright Vision TechnologiesSoftware Development
100% Remote (U.S.)Full-TimeSenior
Salary100,000 - 150,000 USD per year
Apply NowOpens the employer's application page

Job Details

Experience
6+ years
Required Skills
PythonJavaKubernetesGoGrafanaPrometheusCI/CDLinuxDistributed Systems

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or a related technical discipline.
  • Five or more years of SRE, DevOps, or production engineering experience supporting large-scale distributed systems.
  • Strong programming skills in at least one of Python, Go, or Java.
  • Deep, hands-on experience operating Linux at scale, including networking, performance tuning, and systems-level troubleshooting.
  • Production experience operating Kubernetes and container-based workloads.
  • Strong working knowledge of observability tooling such as Prometheus, Grafana, OpenTelemetry, ELK/EFK, or commercial equivalents.
  • Hands-on experience designing and operating CI/CD pipelines for both infrastructure and applications.
  • Solid understanding of distributed system design, including consistency models, partitioning, and failure semantics.
  • Demonstrated experience leading incident response and conducting effective post-incident reviews.
  • Excellent communication and documentation skills.

Responsibilities

  • Ensure availability, performance, and operational excellence of large-scale distributed systems in production.
  • Apply software engineering principles to infrastructure and operations problems.
  • Push the platform toward higher reliability with lower operational toil.
  • Design, automate, and operate complex services.
  • Lead incident response and conduct effective post-incident reviews.
View Full Description & ApplyYou'll be redirected to the employer's site
100,000 - 150,000 USD per year
Apply Now