Site Reliability Engineer

New
G
GiveCampusEducational Fundraising
Remote-first role based in the U.S.Full-TimeMiddle
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
5+ years
Required Skills
AWSKubernetesCI/CDLinuxTerraformGitHub ActionsDatadog

Requirements

  • Approximately 5+ years of related experience in software engineering, infrastructure, systems engineering, SRE, Platform Engineering, or DevOps.
  • Hands-on experience operating production workloads in AWS.
  • Experience building or maintaining infrastructure using Terraform or a similar infrastructure-as-code tool.
  • Experience with New Relic, Datadog, or another modern observability platform.
  • Experience troubleshooting production incidents and participating in an on-call rotation.
  • Experience building or maintaining CI/CD pipelines.
  • Software development or scripting experience, with the ability to read, debug, and make targeted changes to code.
  • Working knowledge of Linux, networking, distributed systems, and relational databases.
  • Ability to manage a well-scoped project with general direction and provide timely updates.
  • Strong written and verbal communication skills and a collaborative approach.

Responsibilities

  • Operate, maintain, and improve production infrastructure in AWS.
  • Build and maintain infrastructure as code using Terraform.
  • Support workloads running on Kubernetes and Amazon EKS.
  • Improve dashboards, alerts, and service-level indicators using New Relic or comparable observability platforms.
  • Investigate production issues, identify root causes, and implement durable fixes.
  • Participate in the shared 24/7 on-call rotation and contribute to effective incident response.
  • Participate in blameless postmortems and complete follow-up actions that reduce the likelihood or impact of repeat incidents.
  • Partner with product engineers to troubleshoot performance and reliability issues throughout the application stack.
  • Automate repetitive operational tasks and identify opportunities to reduce engineering toil.
  • Maintain and improve CI/CD pipelines and deployment workflows.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now