Senior Site Reliability Engineer
New
A
AkamaiAI infrastructure
Workplace type: remote; Locations: Kraków, Na Zjeździe 11, Kraków, Country code: PLFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- PythonKubernetesGoGrafanaPrometheusTerraform
Requirements
- Demonstrate expertise in SRE, infrastructure, or platform engineering.
- Have extensive operational experience managing large-scale distributed systems.
- Demonstrate expertise in Kubernetes and large-scale containerization systems.
- Define SLOs and work with observability tools such as Prometheus, Grafana, and distributed tracing.
- Demonstrate proficiency in Python or Go for automation.
- Have experience with CI/CD pipelines, deployment safety, and infrastructure-as-code such as Terraform.
- Have an interest in or experience with AI/ML infrastructure, model serving, or GPU workloads.
- Resolve issues independently and maintain accountability throughout the process.
- Collaborate effectively with an engineering team unfamiliar with SRE practices.
Responsibilities
- Build and maintain observability for AI workloads, including telemetry, dashboards, alerts, and SLO/SLI tracking.
- Drive improvements when reliability targets are missed.
- Write automation and tooling to reduce operational toil, improve deployment safety, and accelerate incident response.
- Integrate AI workloads into incident management processes, build runbooks, participate in on-call rotations, and conduct blameless post-mortems.
- Build and maintain CI/CD integrations, deployment safety checks, and rollback automation.
- Collaborate with product engineering teams on reliability, architecture decisions, and operational readiness for product releases.
- Contribute to capacity planning, autoscaling configuration, and workload scheduling for AI compute infrastructure.
View Full Description & ApplyYou'll be redirected to the employer's site