Staff DevOps Engineer, Developer Experience (DevEx)
New
J
JobgetherCloud infrastructure
Fully remote position based in the United States.Full-TimeStaff
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 10+ years of overall engineering experience, including meaningful hands-on software development experience in Go, Python, Java, or similar languages. 7+ years of experience designing, building, and operating cloud infrastructure at significant scale.
- Required Skills
- AWSKubernetesCI/CDTerraform
Requirements
- 10+ years of overall engineering experience, including meaningful hands-on software development experience in Go, Python, Java, or similar languages.
- 7+ years of experience designing, building, and operating cloud infrastructure at significant scale.
- Demonstrated technical leadership across multiple teams through architecture decisions, RFCs, platform strategy, and cross-functional influence.
- Deep hands-on experience designing and operating production multi-account, multi-cluster Kubernetes platforms, particularly EKS and/or GKE.
- Strong proficiency with Infrastructure as Code, including Terraform and Crossplane, and organizational-scale GitOps practices such as Flux or Argo CD.
- Extensive experience designing and operating secure, high-velocity CI/CD platforms; GitLab CI/CD experience is preferred.
- Hands-on experience with service mesh technologies such as Istio in multi-cluster environments.
- Strong experience building and operating observability platforms using tools such as Prometheus, Datadog, Grafana, Loki, Mimir, or Thanos, with attention to cost efficiency.
- Working knowledge of container and software supply-chain security, including image scanning, hardened base images, and vulnerability remediation.
- Experience owning or substantially contributing to cloud cost governance and FinOps initiatives.
- Familiarity with AI-assisted and agentic development tools such as Cursor, Claude Code, or Codex, and the ability to establish safe organizational adoption practices.
- Strong understanding of SLIs, SLOs, alerting, and incident response, including leading blameless postmortems.
- Experience supporting data pipeline and batch-processing platforms such as Airflow, EMR, or Dataproc is a plus.
Responsibilities
- Set the technical roadmap and long-term strategy for the Internal Developer Platform.
- Own and evolve the platform as a product, including Backstage catalogs, scaffolder templates, and CI/CD components.
- Author and review RFCs, architecture documents, and technical proposals across engineering teams.
- Design and operate CI/CD orchestration, standardize build and deployment practices, and lead GitOps strategy.
- Own multi-cluster Kubernetes platform architecture and lifecycle across AWS and GCP.
- Establish cloud IAM, networking, security, and Infrastructure-as-Code standards across accounts and environments.
- Lead service mesh architecture and policies, including Istio configuration and traffic controls.
- Define observability strategy and drive FinOps practices, cost visibility, and waste reduction.
- Set standards for secure AI-assisted and agentic development workflows.
- Lead high-severity incident escalations, retrospectives, and organization-wide reliability practices; mentor senior and staff-level engineers.
View Full Description & ApplyYou'll be redirected to the employer's site