Senior Site Reliability Engineer

New
A
AkamaiAI infrastructure
Workplace type: remote; Locations: Kraków, Na Zjeździe 11, Kraków, Country code: PLFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Required Skills
PythonKubernetesGoGrafanaPrometheusTerraform

Requirements

  • Demonstrate expertise in SRE, infrastructure, or platform engineering.
  • Have extensive operational experience managing large-scale distributed systems.
  • Demonstrate expertise in Kubernetes and large-scale containerization systems.
  • Define SLOs and work with observability tools such as Prometheus, Grafana, and distributed tracing.
  • Demonstrate proficiency in Python or Go for automation.
  • Have experience with CI/CD pipelines, deployment safety, and infrastructure-as-code such as Terraform.
  • Have an interest in or experience with AI/ML infrastructure, model serving, or GPU workloads.
  • Resolve issues independently and maintain accountability throughout the process.
  • Collaborate effectively with an engineering team unfamiliar with SRE practices.

Responsibilities

  • Build and maintain observability for AI workloads, including telemetry, dashboards, alerts, and SLO/SLI tracking.
  • Drive improvements when reliability targets are missed.
  • Write automation and tooling to reduce operational toil, improve deployment safety, and accelerate incident response.
  • Integrate AI workloads into incident management processes, build runbooks, participate in on-call rotations, and conduct blameless post-mortems.
  • Build and maintain CI/CD integrations, deployment safety checks, and rollback automation.
  • Collaborate with product engineering teams on reliability, architecture decisions, and operational readiness for product releases.
  • Contribute to capacity planning, autoscaling configuration, and workload scheduling for AI compute infrastructure.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now