Senior Site Reliability Engineer

New
A
AlpacaFinancial Services
Remote - AmericasFull-TimeSenior
SalaryCompetitive Salary & Stock Options
Apply NowOpens the employer's application page

Job Details

Experience
4+ years
Required Skills
PostgreSQLPythonKubernetesGoLinux

Requirements

  • 4+ years of experience in SRE, DevOps, Platform/Infrastructure, or backend engineering with production operations ownership.
  • Hands-on experience operating production services on Kubernetes.
  • Experience shipping infrastructure as code in a GitOps workflow.
  • Solid working knowledge of PostgreSQL in production (query plans, indexing, schema trade-offs, online migrations).
  • Cloud networking fundamentals including VPCs, routing, L4/L7 load balancing, DNS, and TLS.
  • Proficiency with Linux at the operator level.
  • Practiced in structured incident response and debugging.
  • Proficiency in Go or Python.
  • Strong written and verbal communication skills.
  • Genuine interest in growing PostgreSQL and DBA expertise.

Responsibilities

  • Operate production systems including on-call rotation, incident response, and post-mortems.
  • Define and refine SLIs, SLOs, and error budgets for product teams.
  • Enhance observability stack covering metrics, logs, traces, and alerting.
  • Manage infrastructure as code within a GitOps workflow for cloud and Kubernetes resources.
  • Perform PostgreSQL database tasks including performance tuning, schema reviews, online migrations, and HA/DR management.
  • Mentor engineers on reliability and database fundamentals through reviews and pairing.
View Full Description & ApplyYou'll be redirected to the employer's site
Competitive Salary & Stock Options
Apply Now