Senior Site Reliability Engineer (AI/ML Platform)

New
L
Link GroupAI/ML Infrastructure
Warszawa, -, Gdańsk, -, Kraków, -, Wrocław, -, Poznań, -, Lublin, -, Łódź, -, Białystok, -, Olsztyn, -, Szczecin, -, Warszawa, Country code: PLFull-TimeSenior
Salary20,000 - 27,000 PLN per month
Apply NowOpens the employer's application page

Job Details

Required Skills
PythonArtificial IntelligenceKubernetesMachine LearningGoGrafanaPrometheusCI/CDTerraform

Requirements

  • Deep, practical experience managing large-scale containerized environments with Kubernetes.
  • Strong programming skills in Python or Go.
  • Hands-on experience with observability tools like Prometheus, Grafana, and distributed tracing.
  • Experience with Infrastructure-as-Code (IaC) using Terraform.
  • Proven track record of managing complex, large-scale distributed systems.
  • Direct experience or deep interest in AI/ML infrastructure, such as model serving or GPU-accelerated workloads.
  • Ability to take full ownership of technical issues and implement preventative solutions.
  • Excellent collaboration and mentoring skills.

Responsibilities

  • Design and implement comprehensive observability, including robust telemetry, dashboards, and intelligent alerting.
  • Define and track SLOs/SLIs to ensure service reliability.
  • Automate manual toil and create self-healing systems using Python or Go.
  • Lead incident management, participate in on-call rotations, and conduct post-mortems.
  • Develop CI/CD pipelines with automated safety checks and rollback capabilities.
  • Collaborate with product teams to advise on reliability and operational architecture.
View Full Description & ApplyYou'll be redirected to the employer's site
20,000 - 27,000 PLN per month
Apply Now