Sr. Software Engineer, Site Reliability
New
B
BloomerangNonprofit Technology
Within the U.S. and select Canadian Provinces onlyFull-TimeSenior
Salary$114,800 - $150,000
Apply NowOpens the employer's application page
Job Details
- Required Skills
- Node.jsPHPPostgreSQLSQL
Requirements
- Hands-on Site Reliability Engineering experience.
- Experience helping establish or mature SRE practices.
- Strong knowledge of SLIs, SLOs, error budgets, observability, and automation.
- Experience building monitoring and telemetry using tools like Honeycomb, New Relic, Grafana, CloudWatch, or Kibana.
- Experience with production incident management and blameless post-incident reviews.
- Strong programming and scripting skills for automation and operational tooling.
- Strong SQL and relational database skills, with PostgreSQL experience preferred.
- Experience troubleshooting cloud-hosted applications across stacks like PHP, .NET, or Node.js.
- Experience using AI-assisted tools for engineering workflows, troubleshooting, and code analysis.
- Strong debugging skills and ability to navigate unfamiliar codebases.
Responsibilities
- Own complex production support escalations and ticket triage.
- Partner with Software Engineering to investigate complex production issues and drive permanent solutions.
- Lead incident response from triage through root cause analysis and blameless post-incident reviews.
- Build observability across products and services using metrics, logs, traces, and dashboards.
- Define and mature SLIs and SLOs to measure system reliability.
- Develop synthetic monitoring for critical customer journeys.
- Identify and automate recurring operational toil.
- Participate in a rotating on-call schedule.
View Full Description & ApplyYou'll be redirected to the employer's site