Senior Site Reliability Engineer

New
F
FingerprintFraud Detection
Fingerprint is a globally dispersed, 100% remote company. ... people can join our workforce from almost any countryFull-TimeSenior
Salary152,000 - 205,000 USD per year
Apply NowOpens the employer's application page

Job Details

Languages
English
Experience
6–10 years
Required Skills
AWSPythonKubernetesGoGrafanaPrometheusRedisTerraformDatadog

Requirements

  • 6–10 years of experience in SRE, production engineering, infrastructure, or backend engineering.
  • Meaningful time spent responsible for systems in production within cloud-based environments (AWS preferred).
  • Track record of end-to-end system ownership, including design, shipping, and operation.
  • Hands-on experience defining and operating against SLIs, SLOs, and error budgets.
  • Strong incident management skills with experience as a primary responder for customer-facing incidents.
  • Depth in distributed systems failure modes in high-throughput, low-latency environments.
  • Proficiency in cloud infrastructure fundamentals: networking, load balancing, and containerization (EKS/Kubernetes).
  • Strong experience managing infrastructure through code (Terraform).
  • Solid programming skills in Go, Python, or a comparable language.
  • Fluency with observability tooling such as Datadog, Prometheus, Grafana, or OpenTelemetry.
  • Hands-on experience operating Redis/ElastiCache in production (cluster/shard management, scaling strategies).
  • Strong written and verbal communication skills in English.

Responsibilities

  • Own the reliability of core production systems end to end by instrumenting, setting targets, and managing behavior under real traffic.
  • Define and maintain SLIs and SLOs, wiring them into dashboards and alerts to guide engineering priorities.
  • Lead incident response investigations, restore service, and document findings in actionable postmortems.
  • Build secure, resilient, and cost-efficient infrastructure, focusing on failure modes like backpressure, load shedding, and graceful degradation.
  • Conduct capacity and performance analysis using real data, including load testing, profiling, and headroom planning.
  • Manage infrastructure as code primarily using Terraform and design developer-facing tooling to reduce operational toil.
  • Partner with product teams on production readiness, including runbooks, failure mode analysis, and on-call handoffs.
  • Participate in the on-call rotation and improve documentation and escalation clarity to reduce alert fatigue.
View Full Description & ApplyYou'll be redirected to the employer's site
152,000 - 205,000 USD per year
Apply Now