Staff Site Reliability Engineer

New
F
FingerprintFraud Detection
Fingerprint is a globally dispersed, 100% remote company. We are an all-remote company and people can join our workforce from almost any country.Full-TimeStaff
Salary177,000 - 240,000 USD per year
Apply NowOpens the employer's application page

Job Details

Experience
10+ years of engineering experience, with 3+ years as an SRE, production engineer, or reliability-focused Staff engineer
Required Skills
AWSKubernetesTypeScriptGoDatadogDistributed Systems

Requirements

  • 10+ years of engineering experience.
  • 3+ years as an SRE, production engineer, or reliability-focused Staff engineer operating across multiple teams.
  • Proven experience with SLI/SLO design and error budget implementation.
  • Strong incident leadership skills in high-severity, customer-facing environments.
  • Hands-on depth in distributed systems failure modes (cache/database saturation, cascading failure, retry storms).
  • Fluent in Kubernetes, AWS, and modern observability tooling (e.g., Datadog).
  • Proficiency in reading/writing production code (Go, TypeScript, or similar) and infrastructure as code.
  • Demonstrated ability to lead through influence across non-managed teams.
  • Experience coaching engineers to own reliability outcomes.
  • Exceptional written communication skills for async, documented decision-making.
  • Experience using AI tools for incident investigation, analysis, and building operational tooling.

Responsibilities

  • Define SLIs and SLOs for critical request paths (identification, events, server APIs, client agents) and tie them to team decisions.
  • Introduce error budgets to balance reliability investments against feature work.
  • Strengthen incident lifecycles, from detection and communication to postmortem quality and follow-up.
  • Lead reliability reviews for high-risk changes and implement failure testing (game days, chaos exercises).
  • Develop Staff and Lead engineers as reliability leaders within their groups.
  • Codify lightweight production readiness, on-call standards, and runbook practices.
  • Lead AI adoption for incident investigation, observability, and runbook automation.
  • Dig into production during incidents and build tooling, dashboards, or reference implementations.
View Full Description & ApplyYou'll be redirected to the employer's site
177,000 - 240,000 USD per year
Apply Now