Senior Site Reliability Engineer

New
S
SunCore DigitalDigital platforms
Source API remote eligibility restrictions: United Kingdom, United StatesFull-TimeSenior
Salary145,000 - 185,000 USD per year
Apply NowOpens the employer's application page

Job Details

Experience
Five or more years of experience in site reliability, production engineering, infrastructure engineering, DevOps, or a comparable role.
Required Skills
CI/CDRESTful APIsDistributed Systems

Requirements

  • Have five or more years of experience in site reliability, production engineering, infrastructure engineering, DevOps, or a comparable role.
  • Have hands-on experience supporting cloud-hosted applications and distributed systems.
  • Have implemented reliability improvements, not only produced assessments.
  • Have experience with performance, load, stress, or capacity testing.
  • Be able to diagnose application, database, network, and infrastructure bottlenecks.
  • Have strong knowledge of monitoring, logging, metrics, alerting, and incident response.
  • Have experience with reliability automation and scripting.
  • Have experience working with CI/CD systems.
  • Have experience troubleshooting APIs, backend services, databases, queues, caches, and external integrations.
  • Understand timeout, retry, rate-limit, circuit-breaking, and graceful-degradation patterns.
  • Be able to work directly in unfamiliar codebases and environments.
  • Have strong technical documentation and communication skills.

Responsibilities

  • Assess reliability risks, dependencies, failure modes, and single points of failure across applications and infrastructure.
  • Establish performance and capacity-testing programs, baselines, safe operating limits, and early-warning indicators.
  • Design and execute load, stress, endurance, scalability, and failure tests, then repeat testing to validate changes.
  • Implement shared reliability improvements and coordinate application-specific changes with system-owning teams.
  • Assess and improve logging, metrics, tracing, health checks, dashboards, and alerting.
  • Define service-level indicators and initial reliability objectives with system owners and engineering leadership.
  • Create and test incident-response and troubleshooting runbooks, and help establish on-call and escalation practices.
  • Automate reliability work such as diagnostics, incident triage, capacity checks, and evidence collection.
  • Integrate reliability checks and test logic into CI/CD with DevOps and system-owning teams.
  • Document findings, tested capacity, unresolved risks, automation safeguards, and ownership handoffs.
View Full Description & ApplyYou'll be redirected to the employer's site
145,000 - 185,000 USD per year
Apply Now