Site Reliability Engineer
New
G
GiveCampusEducational Fundraising
Remote-first role based in the U.S.Full-TimeMiddle
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- AWSKubernetesCI/CDLinuxTerraformGitHub ActionsDatadog
Requirements
- Approximately 5+ years of related experience in software engineering, infrastructure, systems engineering, SRE, Platform Engineering, or DevOps.
- Hands-on experience operating production workloads in AWS.
- Experience building or maintaining infrastructure using Terraform or a similar infrastructure-as-code tool.
- Experience with New Relic, Datadog, or another modern observability platform.
- Experience troubleshooting production incidents and participating in an on-call rotation.
- Experience building or maintaining CI/CD pipelines.
- Software development or scripting experience, with the ability to read, debug, and make targeted changes to code.
- Working knowledge of Linux, networking, distributed systems, and relational databases.
- Ability to manage a well-scoped project with general direction and provide timely updates.
- Strong written and verbal communication skills and a collaborative approach.
Responsibilities
- Operate, maintain, and improve production infrastructure in AWS.
- Build and maintain infrastructure as code using Terraform.
- Support workloads running on Kubernetes and Amazon EKS.
- Improve dashboards, alerts, and service-level indicators using New Relic or comparable observability platforms.
- Investigate production issues, identify root causes, and implement durable fixes.
- Participate in the shared 24/7 on-call rotation and contribute to effective incident response.
- Participate in blameless postmortems and complete follow-up actions that reduce the likelihood or impact of repeat incidents.
- Partner with product engineers to troubleshoot performance and reliability issues throughout the application stack.
- Automate repetitive operational tasks and identify opportunities to reduce engineering toil.
- Maintain and improve CI/CD pipelines and deployment workflows.
View Full Description & ApplyYou'll be redirected to the employer's site