Staff Site Reliability Engineer
New
V
VeeamCloud SaaS
Remote, United States, Standard shifts, aligned to business hours in a 3×8h global rotation; 8 hour daytime rotationsFull-TimeStaff
SalaryU.S. Geographic Zones & Compensation Ranges (TTC / OTE): Zone 1: San Francisco Bay Area, New York City Boroughs $234,840 — $436,080 USD; Zone 2: Washington, California (excluding San Francisco Bay Area) $215,270 — $399,740 USD; Zone 3: Texas, Illinois, North Carolina, Colorado, Massachusetts, Pennsylvania, Virginia, Oregon, Nevada, Hawaii, New York (excluding NYC boroughs); Sales roles located in Georgia, Ohio, and Arizona $195,700 — $363,400 USD; Zone 4: All other US locations $170,259 — $316,158 USD. The compensation range reflects expected total target compensation (TTC), inclusive of base pay and a competitive performance-based bonus.
Apply NowOpens the employer's application page
Job Details
- Experience
- 8+ years in software engineering for cloud-based products
- Required Skills
- JavaKubernetesC#GoCI/CDTerraformDistributed Systems
Requirements
- Have 8+ years of software engineering experience for cloud-based products.
- Have significant experience designing and operating distributed systems at scale.
- Be proficient in at least one of C#, Java, Go, or TypeScript/Node.js.
- Have experience writing production-grade services and libraries.
- Have hands-on experience with Kubernetes.
- Have hands-on experience with infrastructure as code using Terraform or Pulumi.
- Have hands-on experience with CI/CD, such as GitHub Actions, GitLab, or ArgoCD.
- Bring practical observability expertise across metrics, tracing, and logging.
- Have experience turning SLOs and error budgets into engineering workflows.
- Be able to lead cross-team initiatives, influence architecture, and deliver reliability outcomes.
- Be comfortable with a follow-the-sun on-call model, including 8-hour daytime rotations and coverage.
Responsibilities
- Build reusable reliability libraries, services, and controllers for product teams.
- Define observability data models and implement telemetry pipelines for metrics, logs, and traces.
- Develop SLI/SLO and error-budget policies, and tools that let teams declare SLOs in code and gate releases.
- Build progressive delivery primitives, automated rollback, and release validation services and CI/CD integrations.
- Create fault-injection APIs, chaos experiments, traffic shadowing, and load/performance harnesses.
- Develop Terraform or Pulumi modules, Kubernetes operators, Helm charts, and reference microservice templates.
- Design distributed, multi-region services, initially on Azure, with attention to failure modes, graceful degradation, and operability.
- Instrument systems and automate detection and response while keeping alerts actionable.
- Lead complex daytime incidents, facilitate blameless learning, and implement systemic fixes in code.
- Mentor senior engineers and lead design reviews, architecture decision records, and pair programming.
View Full Description & ApplyYou'll be redirected to the employer's site