Senior Site Reliability Engineer
New
P
PreciselyData Software
100% remote within the U.S., Mountain or Pacific Time zones preferred.Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- AWSPythonBashCI/CDLinuxTerraformAnsibleDatadog
Requirements
- Bachelor's degree in Computer Science, Engineering, or equivalent practical experience.
- 5+ years of systems or infrastructure engineering experience in an enterprise production environment.
- Advanced proficiency with Linux (RHEL/Oracle Linux) in hybrid environments.
- Proficiency with Terraform and experience with Ansible role development.
- Intermediate experience with AWS services: EC2, ECS, S3, VPC, IAM, and CloudWatch.
- Proficiency with Python or Bash for automation.
- Demonstrated experience designing monitoring architecture and alerting strategy, preferably with Datadog.
- Solid understanding of TCP/IP, DNS, load balancing, and distributed systems.
- Experience designing and managing CI/CD pipelines and deployment automation.
- Ability to define SLOs and lead Operational Readiness Reviews.
- Active, proficient use of AI coding assistants (GitHub Copilot, Claude, or equivalent).
Responsibilities
- Define and maintain reliability standards including SLOs, SLIs, and error budgets.
- Build and maintain infrastructure-as-code and observability tooling using Terraform, Ansible, and Datadog.
- Partner with engineering teams to embed reliability, scalability, and recovery considerations into service design.
- Lead complex P1/P2 incident responses, serve as incident commander, and conduct root cause analyses.
- Author and maintain operational runbooks, reliability backlogs, and knowledge resources.
- Utilize AI tools for automation, troubleshooting, runbook creation, and architecture documentation.
- Mentor SREs and coach engineering teams on operational best practices and production diagnostics.
View Full Description & ApplyYou'll be redirected to the employer's site