Site Reliability Engineer
New
A
ArborEducation Technology
United KingdomFull-Time
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- NginxPrometheusTerraformDatadog
Requirements
- Experience in performance monitoring and analysis.
- Capacity planning experience.
- Scripting and automation skills with experience in relevant technologies.
- Experience with Infrastructure as Code, specifically Terraform.
- Understanding of relational database technologies and their cloud versions (e.g., AWS Aurora).
- Experience with messaging and distributed asynchronous workloads.
- Experience with nginx or similar technologies.
- Familiarity with SRE processes.
- Awareness of DevOps principles such as the 3 ways and 5 ideals.
Responsibilities
- Proactively monitor and analyse platform performance.
- Collaborate with engineering teams to address performance bottlenecks and ensure scalability.
- Assist engineering teams with implementing and reviewing SLOs.
- Continually improve observability through monitoring, alerting, and dashboards using tools such as DataDog or Prometheus.
- Ensure the service is highly available and resilient while championing best practices in design.
- Devise runbooks and run game sessions to test disaster recovery plans, high availability, and backups.
- Conduct capacity assessments and plan for scaling to meet business needs.
- Participate in incident response, troubleshooting, and blameless postmortems.
- Develop and maintain technical playbooks and documentation.
View Full Description & ApplyYou'll be redirected to the employer's site