Senior Site Reliability Engineer
New
R
ReplitSoftware Development
Remote - USFull-TimeSenior
Salary$210K - $275K
Apply NowOpens the employer's application page
Job Details
- Experience
- 4-8 years of experience in Site Reliability Engineering or similar roles
- Required Skills
- PythonKubernetesGoDevOpsTerraformAnsibleDistributed Systems
Requirements
- 4-8 years of experience in Site Reliability Engineering or similar roles such as DevOps or Systems Engineering.
- Strong programming skills in languages used for automation such as Python or Go.
- Deep understanding of distributed systems.
- Experience with container orchestration platforms, specifically Kubernetes.
- Proficiency with cloud-native technologies.
- Proven track record of implementing and maintaining monitoring and observability solutions.
- Strong incident management skills with experience leading incident response.
- Experience with infrastructure as code and configuration management tools.
Responsibilities
- Design and implement comprehensive monitoring, alerting, and logging systems using observability tools.
- Architect and implement infrastructure automation using tools like Terraform, Ansible, or Pulumi.
- Design and maintain CI/CD pipelines to enable reliable and consistent deployments.
- Work with product and engineering teams to define and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs).
- Lead incident response efforts, conduct post-mortems, and develop runbooks to reduce MTTR.
- Identify and resolve performance bottlenecks and implement capacity planning strategies.
View Full Description & ApplyYou'll be redirected to the employer's site