Senior Site Reliability Engineer

New
T
Texas Sports AcademyEducation Technology
United States. United Kingdom. Ukraine. MexicoContractSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
Eight or more years of hands-on site reliability, infrastructure, or production engineering work.
Required Skills
AWSGrafanaPrometheusCI/CDDatadog

Requirements

  • Eight or more years of hands-on site reliability, infrastructure, or production engineering experience.
  • Deep expertise in AWS, including networking, IAM, VPCs, and system failure modes.
  • Mastery of monitoring and observability (e.g., Datadog, Grafana, Prometheus).
  • Experience building and maintaining fast, safe, and reversible CI/CD deployment pipelines.
  • Ability to write and maintain production code and configuration.
  • AI-first mindset with proficiency in using AI coding tools for workflows and review.
  • Strong written communication skills for creating actionable engineering audit reports.
  • Reliable internet and a quiet dedicated workspace.
  • Nice-to-have: Experience leading on-call rotations and conducting incident postmortems.
  • Nice-to-have: Proven experience with cloud cost optimization.
  • Nice-to-have: Understanding of the intersection between reliability and security.

Responsibilities

  • Audit systems end to end including infrastructure, deployment pipelines, monitoring, alerting, and incident response.
  • Write clear audit reports detailing findings, impact, and remediation steps for the engineering team.
  • Ship the code and configuration changes necessary to implement audit recommendations.
  • Level up reliability practices including on-call management, incident handling, SLO definition, and postmortems.
  • Advise on architecture for scale to prepare for future growth.
  • Partner with the engineering team through pairing, reviews, and clean hand-offs.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now