Senior Site Reliability Engineer
New
G
Gradle Inc.Software Observability
North America (PST), PSTFull-TimeSenior
Salary$150-190k
Apply NowOpens the employer's application page
Job Details
- Languages
- English
- Experience
- 5+ years
- Required Skills
- AWSPythonBashKubernetesGrafanaPrometheusDevOpsTerraform
Requirements
- 5+ years in SRE, DevOps, or equivalent role operating production services at scale.
- Strong Kubernetes experience in production environments.
- Cloud infrastructure expertise, preferably AWS (EKS, RDS, S3, EC2).
- Proficiency with observability tools (Prometheus, Grafana) and Infrastructure as Code (Terraform).
- Track record of incident management and response.
- Knowledge of SRE best practices (SLAs, SLOs).
- Scripting proficiency (Python, Bash) for automation.
- Experience with 24/7 on-call rotations.
- Strong written and verbal English communication.
Responsibilities
- Operate and maintain all Develocity instances and supporting services.
- Participate in a follow-the-sun on-call rotation, owning incident response and troubleshooting issues across the stack.
- Drive automation across application deployment, upgrades, monitoring, self-healing, and recovery.
- Build and maintain observability for all managed services (logging, metrics, tracing, and alerting).
- Work with engineering teams to build reliability into features from the start.
- Run incident response and retrospectives, and make sure we learn from them.
- Own disaster recovery, backups, and business continuity.
- Communicate with customers during incidents and maintenance windows.
- Optimize performance, resource usage, and costs.
- Help evolve our SaaS operations as we grow.
View Full Description & ApplyYou'll be redirected to the employer's site