Senior Site Reliability Engineer

New
P
PlanetSpace and Data
US, Remote; Canada, RemoteFull-TimeSenior
SalaryUS National Salary Range $142,800 — $178,500 USD
Apply NowOpens the employer's application page

Job Details

Experience
6+ years
Required Skills
PythonBashCloud ComputingKubernetesGrafanaPrometheusCI/CDTerraformAnsible

Requirements

  • 6+ years of experience building services that leverage cloud-native infrastructure and tooling
  • Bachelor’s degree in Computer Science or similar
  • Experience deploying and maintaining bare-metal and cloud kubernetes through tools such as Talos, RKE2, Proxmox, or k3s
  • Proficiency with Terraform, Ansible, Helm, Kustomize, and/or similar IaC / GitOps tooling
  • Experience with CI/CD tooling, such as Jenkins, GitLab CI/CD, Argo CD, or CircleCI
  • Experience successfully building, releasing, and supporting highly available, consistently performant services
  • Knowledge of hardware and network level implications of on-prem compute
  • Experience with platform optimization, particularly resource optimization, management, and cluster tuning in a constrained environment
  • Ability to observe and troubleshoot distributed systems with tools such as Alloy, Prometheus, Grafana, and OpenTelemetry
  • Advanced skills in Python, Bash, and other tooling as appropriate to build services and meet product goals
  • Excellent communication skills and the ability to work through collaboration with cross-functional engineering teams
  • Experience working with Jira for task management and progress tracking

Responsibilities

  • Build and deploy computing services and infrastructure in customer environments for a next-generation satellite operations and image processing end-to-end platform
  • Operate in a high-impact, tight knit team to architect novel systems for air-gapped deployments at scale
  • Clarify and surface requirements from ambiguous use cases defined by cross-functional stakeholders, including internal users and external customers
  • Responsible for operations such as deployments, service orchestration, and documentation for cross platform stakeholders
  • Scale architecture while ensuring availability of services
  • Improve reliability and scalability by resolving edge cases, studying failure modes, and writing tests
  • Participate in on-call rotations to ensure operational excellence
View Full Description & ApplyYou'll be redirected to the employer's site
US National Salary Range $142,800 — $178,500 USD
Apply Now