Staff Site Reliability Engineer

New
R
ReplitSoftware Engineering
Remote - United StatesFull-TimeStaff
Salary$250K - $325K
Apply NowOpens the employer's application page

Job Details

Experience
8-10 years
Required Skills
DockerPythonGCPKubernetesGoTerraformDistributed Systems

Requirements

  • 8-10 years of experience in Site Reliability Engineering, DevOps, or Systems/Infrastructure Engineering.
  • Strong programming proficiency in Python or Go with an emphasis on high-quality, well-tested code.
  • Deep understanding of distributed systems and service-oriented architectures.
  • Extensive experience with container orchestration platforms, specifically Kubernetes.
  • Proven track record of designing and maintaining monitoring and observability solutions (metrics, logging, tracing).
  • Strong incident management and response skills for complex production systems.
  • Proficiency with infrastructure as code tools such as Terraform or Pulumi.
  • Excellent written and verbal communication skills.
  • Strong interpersonal skills with experience mentoring engineers.
  • Willingness to analyze and debug any layer of the technology stack.

Responsibilities

  • Design and lead the implementation of comprehensive monitoring, logging, and tracing solutions.
  • Define, implement, and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs).
  • Act as a senior leader during high-impact incidents and conduct thorough post-mortems.
  • Architect and improve CI/CD pipelines and infrastructure automation using Terraform or Pulumi.
  • Optimize performance and capacity planning for large-scale Kubernetes, Docker, and GCP deployments.
  • Debug complex technical problems across the stack to implement long-term structural fixes.
  • Mentor and educate engineers across all levels to promote reliability as a core company value.
View Full Description & ApplyYou'll be redirected to the employer's site
$250K - $325K
Apply Now