Senior Site Reliability Engineer
New
J
JobgetherCloud Infrastructure
Based in the United StatesFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- AWSPythonBashKubernetesGoCI/CDLinuxTerraform
Requirements
- 5+ years of experience in Site Reliability Engineering, DevOps, platform engineering, cloud infrastructure, or a closely related discipline.
- Strong hands-on experience designing and operating production workloads in AWS, including networking, IAM, compute, storage, databases, DNS, and managed Kubernetes.
- Advanced Terraform expertise, including reusable modules, remote state, dependency management, and infrastructure lifecycle management.
- Strong experience with observability and production telemetry, using technologies such as Prometheus, Grafana, CloudWatch, New Relic, ELK, OpenSearch, or comparable platforms.
- Deep Kubernetes knowledge covering EKS, Helm, workload scheduling, networking, storage, autoscaling, upgrades, and production operations.
- Strong Linux systems expertise, with the ability to diagnose issues involving CPU, memory, disk, networking, processes, and application dependencies.
- Proficiency in Python, Bash, Go, or another general-purpose programming language.
- Experience troubleshooting complex production incidents and contributing to incident response and postmortems.
- Experience designing or operating CI/CD systems such as GitHub Actions, Jenkins, GitLab CI, or Argo CD.
- Working knowledge of cloud security practices, including IAM, encryption, secrets management, and network segmentation.
Responsibilities
- Define, measure, and continuously improve service reliability through SLIs, SLOs, error budgets, availability targets, and capacity planning.
- Build and maintain comprehensive observability across infrastructure and applications, including metrics, logs, traces, dashboards, and actionable alerting.
- Participate in production incident response, troubleshooting, root-cause analysis, and blameless post-incident reviews.
- Identify repetitive operational tasks and replace them with reliable automation using Python, Bash, Go, CI/CD tooling, and platform APIs.
- Design, develop, review, and maintain reusable Terraform modules and infrastructure-as-code patterns for AWS environments.
- Operate and improve Kubernetes platforms, including EKS clusters, workloads, Helm deployments, autoscaling, and networking.
- Partner with development teams to improve application operability, instrumentation, deployment patterns, and production readiness.
View Full Description & ApplyYou'll be redirected to the employer's site