Senior Site Reliability Engineer

New
J
JobgetherCloud Infrastructure
Based in the United StatesFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
5+ years
Required Skills
AWSPythonBashKubernetesGoCI/CDLinuxTerraform

Requirements

  • 5+ years of experience in Site Reliability Engineering, DevOps, platform engineering, cloud infrastructure, or a closely related discipline.
  • Strong hands-on experience designing and operating production workloads in AWS, including networking, IAM, compute, storage, databases, DNS, and managed Kubernetes.
  • Advanced Terraform expertise, including reusable modules, remote state, dependency management, and infrastructure lifecycle management.
  • Strong experience with observability and production telemetry, using technologies such as Prometheus, Grafana, CloudWatch, New Relic, ELK, OpenSearch, or comparable platforms.
  • Deep Kubernetes knowledge covering EKS, Helm, workload scheduling, networking, storage, autoscaling, upgrades, and production operations.
  • Strong Linux systems expertise, with the ability to diagnose issues involving CPU, memory, disk, networking, processes, and application dependencies.
  • Proficiency in Python, Bash, Go, or another general-purpose programming language.
  • Experience troubleshooting complex production incidents and contributing to incident response and postmortems.
  • Experience designing or operating CI/CD systems such as GitHub Actions, Jenkins, GitLab CI, or Argo CD.
  • Working knowledge of cloud security practices, including IAM, encryption, secrets management, and network segmentation.

Responsibilities

  • Define, measure, and continuously improve service reliability through SLIs, SLOs, error budgets, availability targets, and capacity planning.
  • Build and maintain comprehensive observability across infrastructure and applications, including metrics, logs, traces, dashboards, and actionable alerting.
  • Participate in production incident response, troubleshooting, root-cause analysis, and blameless post-incident reviews.
  • Identify repetitive operational tasks and replace them with reliable automation using Python, Bash, Go, CI/CD tooling, and platform APIs.
  • Design, develop, review, and maintain reusable Terraform modules and infrastructure-as-code patterns for AWS environments.
  • Operate and improve Kubernetes platforms, including EKS clusters, workloads, Helm deployments, autoscaling, and networking.
  • Partner with development teams to improve application operability, instrumentation, deployment patterns, and production readiness.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now