Lead Kubernetes Platform Engineer
New
E
EverOpsCloud infrastructure
U.S.-Based Virtual Operating CenterFull-TimeLead
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 8+ years in DevOps, SRE, Platform, or Infrastructure Engineering, including 4+ years operating production Kubernetes
- Required Skills
- PythonKubernetesTerraformHelm
Requirements
- 8+ years in DevOps, SRE, Platform, or Infrastructure Engineering, including 4+ years operating production Kubernetes.
- Prior experience in a technical lead, staff, or principal-level role.
- Deep production experience with Amazon EKS at large scale, such as multiple production clusters, thousands of nodes, or tens of thousands of pods.
- Hands-on ownership of EKS version upgrades across multiple production clusters, including API deprecations, add-on compatibility, node rotation, and rollback planning.
- Working knowledge of the EKS standard and extended support lifecycle.
- Advanced production experience with Karpenter v1+, including NodePools, disruption and consolidation, weighting, and instance-type flexibility.
- Strong understanding of EC2 instance families and generations, CPU architecture differences, network performance limits, and workload bottlenecks.
- Experience migrating production workloads to Graviton or other ARM64 platforms, including multi-architecture container builds and performance validation.
- Experience modeling compute costs and commitments, including Savings Plans, Reserved Instances, and Spot, using real usage data.
- Solid knowledge of AWS VPC CNI, ingress controllers, Gateway API, service mesh, and load balancing on EKS.
- Advanced Terraform proficiency, plus experience with Helm and GitOps tools such as Argo CD or Flux.
- Experience using Datadog, Prometheus, Grafana, or comparable tooling, and scripting with Python, Go, or Bash.
Responsibilities
- Baseline the multi-cluster production EKS estate, including topology, workload placement, ownership, costs, and Kubernetes version and support status.
- Analyze failure domains, isolation boundaries, and dependency concentration, then design a multi-cluster target architecture.
- Design and implement a repeatable, largely automated EKS upgrade approach, including node rotation, add-on compatibility, and API deprecation management.
- Lead instance sizing and workload-fit analysis across compute families, generations, and node sizes.
- Plan and drive a phased ARM64 migration, including multi-architecture builds, native dependency remediation, and rightsizing against latency SLOs.
- Configure Karpenter NodePools, weights, and node overlays, and design Spot patterns for workload requirements.
- Model compute, support, and commitment economics and quantify savings by cost lever.
- Assess and improve ingress, service mesh, CNI, GitOps, and infrastructure-as-code posture across the estate.
- Measure maintenance and toil, set reduction targets, and build automation to reduce operational work.
- Lead the TechPod, coordinate with the customer’s infrastructure owner and AWS specialists, and present plans and recommendations to engineering leadership.
View Full Description & ApplyYou'll be redirected to the employer's site