Senior Site Reliability Engineer
New
S
SunCore DigitalDigital platforms
Source API remote eligibility restrictions: United Kingdom, United StatesFull-TimeSenior
Salary145,000 - 185,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- Five or more years of experience in site reliability, production engineering, infrastructure engineering, DevOps, or a comparable role.
- Required Skills
- CI/CDRESTful APIsDistributed Systems
Requirements
- Have five or more years of experience in site reliability, production engineering, infrastructure engineering, DevOps, or a comparable role.
- Have hands-on experience supporting cloud-hosted applications and distributed systems.
- Have implemented reliability improvements, not only produced assessments.
- Have experience with performance, load, stress, or capacity testing.
- Be able to diagnose application, database, network, and infrastructure bottlenecks.
- Have strong knowledge of monitoring, logging, metrics, alerting, and incident response.
- Have experience with reliability automation and scripting.
- Have experience working with CI/CD systems.
- Have experience troubleshooting APIs, backend services, databases, queues, caches, and external integrations.
- Understand timeout, retry, rate-limit, circuit-breaking, and graceful-degradation patterns.
- Be able to work directly in unfamiliar codebases and environments.
- Have strong technical documentation and communication skills.
Responsibilities
- Assess reliability risks, dependencies, failure modes, and single points of failure across applications and infrastructure.
- Establish performance and capacity-testing programs, baselines, safe operating limits, and early-warning indicators.
- Design and execute load, stress, endurance, scalability, and failure tests, then repeat testing to validate changes.
- Implement shared reliability improvements and coordinate application-specific changes with system-owning teams.
- Assess and improve logging, metrics, tracing, health checks, dashboards, and alerting.
- Define service-level indicators and initial reliability objectives with system owners and engineering leadership.
- Create and test incident-response and troubleshooting runbooks, and help establish on-call and escalation practices.
- Automate reliability work such as diagnostics, incident triage, capacity checks, and evidence collection.
- Integrate reliability checks and test logic into CI/CD with DevOps and system-owning teams.
- Document findings, tested capacity, unresolved risks, automation safeguards, and ownership handoffs.
View Full Description & ApplyYou'll be redirected to the employer's site