Senior Site Reliability Engineer

New
G
Gradle Inc.Software Observability
North America (PST), PSTFull-TimeSenior
Salary$150-190k
Apply NowOpens the employer's application page

Job Details

Languages
English
Experience
5+ years
Required Skills
AWSPythonBashKubernetesGrafanaPrometheusDevOpsTerraform

Requirements

  • 5+ years in SRE, DevOps, or equivalent role operating production services at scale.
  • Strong Kubernetes experience in production environments.
  • Cloud infrastructure expertise, preferably AWS (EKS, RDS, S3, EC2).
  • Proficiency with observability tools (Prometheus, Grafana) and Infrastructure as Code (Terraform).
  • Track record of incident management and response.
  • Knowledge of SRE best practices (SLAs, SLOs).
  • Scripting proficiency (Python, Bash) for automation.
  • Experience with 24/7 on-call rotations.
  • Strong written and verbal English communication.

Responsibilities

  • Operate and maintain all Develocity instances and supporting services.
  • Participate in a follow-the-sun on-call rotation, owning incident response and troubleshooting issues across the stack.
  • Drive automation across application deployment, upgrades, monitoring, self-healing, and recovery.
  • Build and maintain observability for all managed services (logging, metrics, tracing, and alerting).
  • Work with engineering teams to build reliability into features from the start.
  • Run incident response and retrospectives, and make sure we learn from them.
  • Own disaster recovery, backups, and business continuity.
  • Communicate with customers during incidents and maintenance windows.
  • Optimize performance, resource usage, and costs.
  • Help evolve our SaaS operations as we grow.
View Full Description & ApplyYou'll be redirected to the employer's site
$150-190k
Apply Now