Senior Platform Engineer

New
B
BuildkiteCI/CD platform
Location: ANZ Region; it does need to be located in either the ANZ or PST timezone., ANZ or PST timezoneFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Required Skills
AWSKubernetesCI/CDTerraformDistributed Systems

Requirements

  • Have production AWS, Kubernetes, and infrastructure-as-code experience.
  • Have managed production infrastructure with Terraform or an equivalent tool, including DNS, IAM, monitoring, and secrets.
  • Have experience owning production systems through growth, failures, migrations, and incidents.
  • Have experience building CI/CD, internal-platform, or self-service capabilities for other engineers.
  • Apply strong reliability practices, including observability, incident response, recovery testing, and operational improvement.
  • Have experience building platforms or self-service capabilities that enable secure, reliable delivery.
  • Bring depth in several valued areas, sound judgement, and ability to learn the rest.
  • Useful additional experience includes multi-region or multi-cluster systems, traffic routing, and disaster recovery.
  • Useful additional experience includes progressive delivery, deployment guardrails, and rollback.
  • Useful additional experience includes relational databases, caches, or queues, including migrations, replication, failover, and workload isolation.

Responsibilities

  • Own substantial parts of Buildkite’s AWS and multi-cluster Kubernetes platform from design through production operation.
  • Design reusable, self-service automation for provisioning and operating secure, consistent, and isolated platform environments.
  • Deliver regional architecture, request routing, disaster recovery, and tested failover capabilities.
  • Improve Kubernetes networking, upgrades, patching, scaling, observability, capacity management, and recovery.
  • Make infrastructure and deployments safer through secure defaults, progressive delivery, automated guardrails, clear health signals, and fast rollback.
  • Partner with product teams to diagnose distributed-system problems across service, datastore, network, and platform boundaries.
  • Turn incidents, capacity limits, and operational signals into durable engineering improvements.
  • Write well-tested, observable, documented, and operable systems; lead design discussions, review code, share context, and mentor other engineers.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now