Senior Site Reliability Engineer
New
A
AlpacaFinancial Services
Remote - AmericasFull-TimeSenior
SalaryCompetitive Salary & Stock Options
Apply NowOpens the employer's application page
Job Details
- Experience
- 4+ years
- Required Skills
- PostgreSQLPythonKubernetesGoLinux
Requirements
- 4+ years of experience in SRE, DevOps, Platform/Infrastructure, or backend engineering with production operations ownership.
- Hands-on experience operating production services on Kubernetes.
- Experience shipping infrastructure as code in a GitOps workflow.
- Solid working knowledge of PostgreSQL in production (query plans, indexing, schema trade-offs, online migrations).
- Cloud networking fundamentals including VPCs, routing, L4/L7 load balancing, DNS, and TLS.
- Proficiency with Linux at the operator level.
- Practiced in structured incident response and debugging.
- Proficiency in Go or Python.
- Strong written and verbal communication skills.
- Genuine interest in growing PostgreSQL and DBA expertise.
Responsibilities
- Operate production systems including on-call rotation, incident response, and post-mortems.
- Define and refine SLIs, SLOs, and error budgets for product teams.
- Enhance observability stack covering metrics, logs, traces, and alerting.
- Manage infrastructure as code within a GitOps workflow for cloud and Kubernetes resources.
- Perform PostgreSQL database tasks including performance tuning, schema reviews, online migrations, and HA/DR management.
- Mentor engineers on reliability and database fundamentals through reviews and pairing.
View Full Description & ApplyYou'll be redirected to the employer's site