Staff Site Reliability Engineer - Cloud Platform

New
B
Butterfly NetworkMedical Imaging
This role is eligible to be fully remote in the US, based out of these states: (Please note for all roles based outside of our Burlington, MA or NY, NY offices we are only able to employ in the following states: AL, AZ, CA, CO, CT, DE, FL, GA, IL, IN, KS, KY, ,LA, MA, MD, ME, MI, MN, MO, MT, NC, NE, NH, NJ, NM, NY, OH, OK, PA, SC, TN, TX, UT, VA, VT, WA, or WI.)Full-TimeStaff
Salary$190,000 - $210,000 + bonus + equity + benefits
Apply NowOpens the employer's application page

Job Details

Experience
8+ years
Required Skills
AWSKubernetesDatadog

Requirements

  • 8+ years of experience managing production systems
  • Deep, hands-on production experience with AWS
  • Deep, hands-on production experience with Kubernetes
  • Strong programming/scripting skills and experience building operational automation
  • Proven ownership of a full stack observability platform (e.g., NewRelic, Datadog)
  • Track record of leading incident response and delivering reliability improvements
  • Clear written and verbal communication skills
  • Strong desire to mentor others

Responsibilities

  • Own observability end to end by evolving metrics, logs, and distributed-tracing strategies and building telemetry standards.
  • Drive reliability by establishing service-level objectives (SLOs) and using error budgets to balance velocity and stability.
  • Lead incident response and manage on-call rotations, blameless postmortems, and toil reduction.
  • Operate and improve Kubernetes/EKS workloads on AWS using automation and tooling.
  • Mentor engineers, set operational standards, and influence cross-team architecture decisions.
View Full Description & ApplyYou'll be redirected to the employer's site
$190,000 - $210,000 + bonus + equity + benefits
Apply Now