Platform Reliability Engineer
New
B
Bright Vision TechnologiesSoftware Development
100% Remote (U.S.)Full-TimeSenior
Salary100,000 - 150,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 6+ years
- Required Skills
- PythonJavaKubernetesGoGrafanaPrometheusCI/CDLinuxDistributed Systems
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related technical discipline.
- Five or more years of SRE, DevOps, or production engineering experience supporting large-scale distributed systems.
- Strong programming skills in at least one of Python, Go, or Java.
- Deep, hands-on experience operating Linux at scale, including networking, performance tuning, and systems-level troubleshooting.
- Production experience operating Kubernetes and container-based workloads.
- Strong working knowledge of observability tooling such as Prometheus, Grafana, OpenTelemetry, ELK/EFK, or commercial equivalents.
- Hands-on experience designing and operating CI/CD pipelines for both infrastructure and applications.
- Solid understanding of distributed system design, including consistency models, partitioning, and failure semantics.
- Demonstrated experience leading incident response and conducting effective post-incident reviews.
- Excellent communication and documentation skills.
Responsibilities
- Ensure availability, performance, and operational excellence of large-scale distributed systems in production.
- Apply software engineering principles to infrastructure and operations problems.
- Push the platform toward higher reliability with lower operational toil.
- Design, automate, and operate complex services.
- Lead incident response and conduct effective post-incident reviews.
View Full Description & ApplyYou'll be redirected to the employer's site