Staff Site Reliability Engineer
New
R
ReplitSoftware Engineering
Remote - United StatesFull-TimeStaff
Salary$250K - $325K
Apply NowOpens the employer's application page
Job Details
- Experience
- 8-10 years
- Required Skills
- DockerPythonGCPKubernetesGoTerraformDistributed Systems
Requirements
- 8-10 years of experience in Site Reliability Engineering, DevOps, or Systems/Infrastructure Engineering.
- Strong programming proficiency in Python or Go with an emphasis on high-quality, well-tested code.
- Deep understanding of distributed systems and service-oriented architectures.
- Extensive experience with container orchestration platforms, specifically Kubernetes.
- Proven track record of designing and maintaining monitoring and observability solutions (metrics, logging, tracing).
- Strong incident management and response skills for complex production systems.
- Proficiency with infrastructure as code tools such as Terraform or Pulumi.
- Excellent written and verbal communication skills.
- Strong interpersonal skills with experience mentoring engineers.
- Willingness to analyze and debug any layer of the technology stack.
Responsibilities
- Design and lead the implementation of comprehensive monitoring, logging, and tracing solutions.
- Define, implement, and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs).
- Act as a senior leader during high-impact incidents and conduct thorough post-mortems.
- Architect and improve CI/CD pipelines and infrastructure automation using Terraform or Pulumi.
- Optimize performance and capacity planning for large-scale Kubernetes, Docker, and GCP deployments.
- Debug complex technical problems across the stack to implement long-term structural fixes.
- Mentor and educate engineers across all levels to promote reliability as a core company value.
View Full Description & ApplyYou'll be redirected to the employer's site