Senior Site Reliability Engineer

New
W
Wikimedia FoundationTechnology, Non-profit
USA, Europe; please note that we are currently able to hire in the following: US States: Arizona, California, Colorado, Connecticut, District of Columbia*, Florida, Georgia, Idaho, Illinois, Indiana, Iowa, Maryland, Massachusetts, Michigan, Minnesota, Missouri, New Jersey, New Mexico, New York, North Carolina, Ohio, Oklahoma, Oregon, Pennsylvania, Puerto Rico*, Rhode Island, Tennessee, Texas, Utah, Vermont, Virginia, Washington, West Virginia, Wisconsin and Wyoming (*US Territory or Federal District). Countries: Brazil, Canada, Colombia, France, Germany, Ghana, India, Indonesia, Italy, Kenya*, Mexico, Morocco, Netherlands, Poland, Singapore*, South Africa, Spain, Switzerland and the United Kingdom., Global and asynchronously communicating team.Full-TimeSenior
Salary116,633 - 181,243 USD per year
Apply NowOpens the employer's application page

Job Details

Required Skills
AWSPythonGCPAzureGoPrometheusTerraformAnsibleGitLab

Requirements

  • Experience with Infrastructure as Code and automation tools (e.g., Terraform, Ansible).
  • Proficiency in at least one programming language (e.g., Python, Go, or similar).
  • Experience designing, operating, and optimizing cloud-based systems (AWS, Azure, or GCP).
  • Experience building and maintaining CI/CD pipelines and GitOps workflows (e.g., GitLab, ArgoCD).
  • Familiarity with progressive delivery approaches like canary and blue-green deployments.
  • Experience with incident response, on-call practices, and leading postmortems.
  • Understanding of SRE best practices including SLOs, SLIs, and error budgets.
  • Experience with observability tools (e.g., Prometheus, OpenTelemetry).
  • Strong documentation and communication skills.
  • Ability to work effectively in a distributed, cross-functional environment.

Responsibilities

  • Define, track, and improve Service Level Objectives (SLOs), SLIs, and error budgets.
  • Build and enhance observability systems including metrics, logs, and distributed tracing.
  • Drive reliability engineering practices like capacity planning, load testing, and chaos testing.
  • Improve developer experience by enabling self-service infrastructure and streamlining deployment workflows.
  • Design, implement, and optimize CI/CD and GitOps workflows using GitLab and ArgoCD.
  • Implement secure-by-default infrastructure including IAM, secrets management, and encryption.
  • Continuously optimize infrastructure cost and efficiency using FinOps principles.
  • Establish and track operational metrics such as MTTR, MTTD, and incident frequency.
  • Collaborate with a global and asynchronously communicating team.
View Full Description & ApplyYou'll be redirected to the employer's site
116,633 - 181,243 USD per year
Apply Now