Site Reliability Engineer

New
A
ArborEducation Technology
United KingdomFull-Time
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Required Skills
NginxPrometheusTerraformDatadog

Requirements

  • Experience in performance monitoring and analysis.
  • Capacity planning experience.
  • Scripting and automation skills with experience in relevant technologies.
  • Experience with Infrastructure as Code, specifically Terraform.
  • Understanding of relational database technologies and their cloud versions (e.g., AWS Aurora).
  • Experience with messaging and distributed asynchronous workloads.
  • Experience with nginx or similar technologies.
  • Familiarity with SRE processes.
  • Awareness of DevOps principles such as the 3 ways and 5 ideals.

Responsibilities

  • Proactively monitor and analyse platform performance.
  • Collaborate with engineering teams to address performance bottlenecks and ensure scalability.
  • Assist engineering teams with implementing and reviewing SLOs.
  • Continually improve observability through monitoring, alerting, and dashboards using tools such as DataDog or Prometheus.
  • Ensure the service is highly available and resilient while championing best practices in design.
  • Devise runbooks and run game sessions to test disaster recovery plans, high availability, and backups.
  • Conduct capacity assessments and plan for scaling to meet business needs.
  • Participate in incident response, troubleshooting, and blameless postmortems.
  • Develop and maintain technical playbooks and documentation.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now