Principal Site Reliability Engineer, Platform

New
B
Blue River TechnologyRobotics, Automation
Remote in the United States.Full-TimePrincipal
Salary$174,000 - $305,000/year
Apply NowOpens the employer's application page

Job Details

Experience
Min. of 8 years
Required Skills
AWSPythonJavascriptKubernetesGoRustCI/CDTerraformGitHub Actions

Requirements

  • Minimum of 8 years of experience building and maintaining infrastructure for data-intensive, high-availability applications.
  • Minimum of 6 years of experience building and maintaining public cloud solutions.
  • Deep understanding of cloud orchestration tools such as Kubernetes and Terraform.
  • Deep understanding of software design methodologies, information systems architecture, object-oriented design, and software design patterns.
  • Deep understanding of securing cloud infrastructure (preferably AWS and Kubernetes).
  • Deep experience in one or more of: Golang, Python, JavaScript, Rust.
  • Deep experience in CI/CD tooling (GitHub Actions, ArgoCD, ArgoCD Image Updater, Artifactory).

Responsibilities

  • Architect, scale, and own essential infrastructure.
  • Build and maintain a Kubernetes-based platform supporting multiple teams and services.
  • Build backend services (Golang) to support autonomous systems.
  • Partner with product teams to launch new products on the platform.
  • Grow high availability infrastructure while maintaining key uptime metrics.
  • Build tooling to support platform and development teams.
  • Perform end-to-end performance analysis and implement solutions.
  • Participate in on-call rotation, triaging, and resolving production incidents with root cause analysis.
  • Design and maintain observability infrastructure, dashboards, and log aggregation.
  • Collaborate with security on risk assessments, maintain the risk register, and develop mitigation plans.
View Full Description & ApplyYou'll be redirected to the employer's site
$174,000 - $305,000/year
Apply Now