Senior/Staff Platform Engineer
New
J
JobgetherPlatform engineering
Listing location: US; based in United States, Core hours aligned with Pacific Time, from 8:00 AM to 5:00 PM PST; on-call rotation approximately every four to five weeksFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 10+ years of professional experience ... is preferred for Staff-level candidates
- Required Skills
- PythonJavaKubernetesGoGrafanaPrometheusLinuxTerraformAnsible
Requirements
- Demonstrated experience building and operating production Kubernetes platforms, not only deploying applications onto existing clusters.
- Production programming experience in Go, Python, or Java, including reading, debugging, and maintaining production code and automation.
- Strong experience with production reliability, incident response, troubleshooting, root-cause analysis, and operational ownership.
- Ability to own complex technical initiatives from ambiguous starting points through design, implementation, and production operation.
- Deep Kubernetes infrastructure knowledge, including cluster architecture, networking and CNI, NetworkPolicy, scheduling, resource management, security, and RBAC.
- Strong Linux fundamentals and experience troubleshooting production systems; understanding of DNS, routing, load balancing, connectivity, and cloud/Kubernetes networking.
- Production experience with at least one major cloud platform, such as AWS, GCP, or Alicloud, and infrastructure-as-code experience with Terraform or equivalent.
- Experience with configuration management and automation technologies such as Ansible, Puppet, or similar platforms.
- Observability experience with metrics, logs, traces, dashboards, and alerting platforms such as Prometheus, Grafana, OpenTelemetry, or Datadog.
- Experience with CI/CD infrastructure, Docker, high availability, capacity planning, disaster recovery, and production resilience.
- Availability during core hours aligned with Pacific Time, from 8:00 AM to 5:00 PM PST, and participation in an on-call rotation approximately every four to five weeks.
- 10+ years of relevant professional experience is preferred for Staff-level candidates; a related degree is preferred, with equivalent practical experience considered.
Responsibilities
- Design, build, operate, and improve production Kubernetes platforms, including cluster architecture, networking, workload isolation, security, upgrades, scaling, and reliability.
- Troubleshoot Kubernetes and infrastructure issues across CNI, networking, scheduling, nodes, resource constraints, controllers, and cluster failures.
- Operate highly available infrastructure across cloud, hybrid, virtualized, and bare-metal environments.
- Develop production tooling and automation using Go, Python, or Java to improve platform operations, deployment, troubleshooting, and developer experience.
- Own reliability and operational health, contributing to incident response, root-cause analysis, durable remediation, and disaster recovery.
- Define and improve SLOs, SLIs, alerting, and operational processes using logs, metrics, traces, and profiling tools.
- Build and maintain infrastructure as code using Terraform or equivalent technologies, and improve CI/CD and deployment workflows.
- Plan and execute cloud or infrastructure migrations, including dependency analysis, cutover planning, rollback strategies, and production validation.
- Work with customers, engineers, and technical stakeholders to investigate issues, communicate trade-offs, and drive technical solutions.
- Own complex infrastructure initiatives from problem definition through design, implementation, and production operation; contribute to architecture discussions and mentor engineers.
View Full Description & ApplyYou'll be redirected to the employer's site