Principal Site Reliability Engineer, Platform
New
B
Blue River TechnologyRobotics, Automation
Remote in the United States.Full-TimePrincipal
Salary$174,000 - $305,000/year
Apply NowOpens the employer's application page
Job Details
- Experience
- Min. of 8 years
- Required Skills
- AWSPythonJavascriptKubernetesGoRustCI/CDTerraformGitHub Actions
Requirements
- Minimum of 8 years of experience building and maintaining infrastructure for data-intensive, high-availability applications.
- Minimum of 6 years of experience building and maintaining public cloud solutions.
- Deep understanding of cloud orchestration tools such as Kubernetes and Terraform.
- Deep understanding of software design methodologies, information systems architecture, object-oriented design, and software design patterns.
- Deep understanding of securing cloud infrastructure (preferably AWS and Kubernetes).
- Deep experience in one or more of: Golang, Python, JavaScript, Rust.
- Deep experience in CI/CD tooling (GitHub Actions, ArgoCD, ArgoCD Image Updater, Artifactory).
Responsibilities
- Architect, scale, and own essential infrastructure.
- Build and maintain a Kubernetes-based platform supporting multiple teams and services.
- Build backend services (Golang) to support autonomous systems.
- Partner with product teams to launch new products on the platform.
- Grow high availability infrastructure while maintaining key uptime metrics.
- Build tooling to support platform and development teams.
- Perform end-to-end performance analysis and implement solutions.
- Participate in on-call rotation, triaging, and resolving production incidents with root cause analysis.
- Design and maintain observability infrastructure, dashboards, and log aggregation.
- Collaborate with security on risk assessments, maintain the risk register, and develop mitigation plans.
View Full Description & ApplyYou'll be redirected to the employer's site