Senior Site Reliability Engineer

New
R
ReplitSoftware Development
Remote - USFull-TimeSenior
Salary$210K - $275K
Apply NowOpens the employer's application page

Job Details

Experience
4-8 years of experience in Site Reliability Engineering or similar roles
Required Skills
PythonKubernetesGoDevOpsTerraformAnsibleDistributed Systems

Requirements

  • 4-8 years of experience in Site Reliability Engineering or similar roles such as DevOps or Systems Engineering.
  • Strong programming skills in languages used for automation such as Python or Go.
  • Deep understanding of distributed systems.
  • Experience with container orchestration platforms, specifically Kubernetes.
  • Proficiency with cloud-native technologies.
  • Proven track record of implementing and maintaining monitoring and observability solutions.
  • Strong incident management skills with experience leading incident response.
  • Experience with infrastructure as code and configuration management tools.

Responsibilities

  • Design and implement comprehensive monitoring, alerting, and logging systems using observability tools.
  • Architect and implement infrastructure automation using tools like Terraform, Ansible, or Pulumi.
  • Design and maintain CI/CD pipelines to enable reliable and consistent deployments.
  • Work with product and engineering teams to define and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs).
  • Lead incident response efforts, conduct post-mortems, and develop runbooks to reduce MTTR.
  • Identify and resolve performance bottlenecks and implement capacity planning strategies.
View Full Description & ApplyYou'll be redirected to the employer's site
$210K - $275K
Apply Now