Lead Kubernetes Platform Engineer

New
E
EverOpsCloud infrastructure
U.S.-Based Virtual Operating CenterFull-TimeLead
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
8+ years in DevOps, SRE, Platform, or Infrastructure Engineering, including 4+ years operating production Kubernetes
Required Skills
PythonKubernetesTerraformHelm

Requirements

  • 8+ years in DevOps, SRE, Platform, or Infrastructure Engineering, including 4+ years operating production Kubernetes.
  • Prior experience in a technical lead, staff, or principal-level role.
  • Deep production experience with Amazon EKS at large scale, such as multiple production clusters, thousands of nodes, or tens of thousands of pods.
  • Hands-on ownership of EKS version upgrades across multiple production clusters, including API deprecations, add-on compatibility, node rotation, and rollback planning.
  • Working knowledge of the EKS standard and extended support lifecycle.
  • Advanced production experience with Karpenter v1+, including NodePools, disruption and consolidation, weighting, and instance-type flexibility.
  • Strong understanding of EC2 instance families and generations, CPU architecture differences, network performance limits, and workload bottlenecks.
  • Experience migrating production workloads to Graviton or other ARM64 platforms, including multi-architecture container builds and performance validation.
  • Experience modeling compute costs and commitments, including Savings Plans, Reserved Instances, and Spot, using real usage data.
  • Solid knowledge of AWS VPC CNI, ingress controllers, Gateway API, service mesh, and load balancing on EKS.
  • Advanced Terraform proficiency, plus experience with Helm and GitOps tools such as Argo CD or Flux.
  • Experience using Datadog, Prometheus, Grafana, or comparable tooling, and scripting with Python, Go, or Bash.

Responsibilities

  • Baseline the multi-cluster production EKS estate, including topology, workload placement, ownership, costs, and Kubernetes version and support status.
  • Analyze failure domains, isolation boundaries, and dependency concentration, then design a multi-cluster target architecture.
  • Design and implement a repeatable, largely automated EKS upgrade approach, including node rotation, add-on compatibility, and API deprecation management.
  • Lead instance sizing and workload-fit analysis across compute families, generations, and node sizes.
  • Plan and drive a phased ARM64 migration, including multi-architecture builds, native dependency remediation, and rightsizing against latency SLOs.
  • Configure Karpenter NodePools, weights, and node overlays, and design Spot patterns for workload requirements.
  • Model compute, support, and commitment economics and quantify savings by cost lever.
  • Assess and improve ingress, service mesh, CNI, GitOps, and infrastructure-as-code posture across the estate.
  • Measure maintenance and toil, set reduction targets, and build automation to reduce operational work.
  • Lead the TechPod, coordinate with the customer’s infrastructure owner and AWS specialists, and present plans and recommendations to engineering leadership.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now