Staff Site Reliability Engineer - Cloud Platform
New
B
Butterfly NetworkMedical Imaging
This role is eligible to be fully remote in the US, based out of these states: (Please note for all roles based outside of our Burlington, MA or NY, NY offices we are only able to employ in the following states: AL, AZ, CA, CO, CT, DE, FL, GA, IL, IN, KS, KY, ,LA, MA, MD, ME, MI, MN, MO, MT, NC, NE, NH, NJ, NM, NY, OH, OK, PA, SC, TN, TX, UT, VA, VT, WA, or WI.)Full-TimeStaff
Salary$190,000 - $210,000 + bonus + equity + benefits
Apply NowOpens the employer's application page
Job Details
- Experience
- 8+ years
- Required Skills
- AWSKubernetesDatadog
Requirements
- 8+ years of experience managing production systems
- Deep, hands-on production experience with AWS
- Deep, hands-on production experience with Kubernetes
- Strong programming/scripting skills and experience building operational automation
- Proven ownership of a full stack observability platform (e.g., NewRelic, Datadog)
- Track record of leading incident response and delivering reliability improvements
- Clear written and verbal communication skills
- Strong desire to mentor others
Responsibilities
- Own observability end to end by evolving metrics, logs, and distributed-tracing strategies and building telemetry standards.
- Drive reliability by establishing service-level objectives (SLOs) and using error budgets to balance velocity and stability.
- Lead incident response and manage on-call rotations, blameless postmortems, and toil reduction.
- Operate and improve Kubernetes/EKS workloads on AWS using automation and tooling.
- Mentor engineers, set operational standards, and influence cross-team architecture decisions.
View Full Description & ApplyYou'll be redirected to the employer's site