Engineering Manager, Site Reliability
I
InstacartGrocery technology
United States - RemoteFull-TimeManager
SalaryCA, NY, CT, NJ $214,000 — $226,000 USD; WA $205,000 — $216,500 USD; OR, DE, ME, MA, MD, NH, RI, VT, DC, PA, VA, CO, TX, IL, HI $196,000 — $207,000 USD; All other states $178,000 — $188,000 USD
Apply NowOpens the employer's application page
Job Details
- Experience
- Seven or more years of experience in software engineering, infrastructure engineering, Site Reliability Engineering, or a related field; two or more years of experience managing, mentoring, or leading engineering teams.
- Required Skills
- Distributed SystemsNetworking
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
- Seven or more years of experience in software engineering, infrastructure engineering, Site Reliability Engineering, or a related field.
- Two or more years of experience managing, mentoring, or leading engineering teams.
- Professional experience with cloud infrastructure, distributed systems, networking, containers, orchestration platforms, or related production technologies.
- Experience leading or participating in production incident response, post-incident reviews, reliability improvement initiatives, and operational readiness practices.
- Experience communicating technical risks, priorities, and tradeoffs to engineering leaders and cross-functional stakeholders.
- Comfort working with ambiguity and engaging directly with operational and organizational problems.
- Preferred experience leading Site Reliability Engineering, platform engineering, infrastructure engineering, or developer productivity teams.
- Preferred experience operating highly available services at scale and improving service-level objectives, observability, capacity, or disaster recovery.
- Preferred experience with infrastructure as code, continuous delivery, monitoring, logging, tracing, and automated remediation.
Responsibilities
- Lead, mentor, and develop a team of Site Reliability Engineers, setting goals, providing feedback, and supporting career growth.
- Set technical direction and priorities for improving system reliability, scalability, availability, performance, and operational readiness.
- Partner with engineering, product, security, infrastructure, and other teams to define reliability standards and influence system design.
- Drive incident response, post-incident learning, service-level objectives, capacity planning, observability, and risk reduction.
- Promote automation and self-service tooling to reduce operational toil and improve deployment confidence.
- Balance near-term operational needs with long-term investments and make tradeoffs as priorities change.
- Communicate reliability risks, project status, tradeoffs, and decisions to technical and non-technical stakeholders.
View Full Description & ApplyYou'll be redirected to the employer's site