Senior Site Reliability Engineer (AI/ML Platform)
New
L
Link GroupAI/ML Infrastructure
Warszawa, -, Gdańsk, -, Kraków, -, Wrocław, -, Poznań, -, Lublin, -, Łódź, -, Białystok, -, Olsztyn, -, Szczecin, -, Warszawa, Country code: PLFull-TimeSenior
Salary20,000 - 27,000 PLN per month
Apply NowOpens the employer's application page
Job Details
- Required Skills
- PythonArtificial IntelligenceKubernetesMachine LearningGoGrafanaPrometheusCI/CDTerraform
Requirements
- Deep, practical experience managing large-scale containerized environments with Kubernetes.
- Strong programming skills in Python or Go.
- Hands-on experience with observability tools like Prometheus, Grafana, and distributed tracing.
- Experience with Infrastructure-as-Code (IaC) using Terraform.
- Proven track record of managing complex, large-scale distributed systems.
- Direct experience or deep interest in AI/ML infrastructure, such as model serving or GPU-accelerated workloads.
- Ability to take full ownership of technical issues and implement preventative solutions.
- Excellent collaboration and mentoring skills.
Responsibilities
- Design and implement comprehensive observability, including robust telemetry, dashboards, and intelligent alerting.
- Define and track SLOs/SLIs to ensure service reliability.
- Automate manual toil and create self-healing systems using Python or Go.
- Lead incident management, participate in on-call rotations, and conduct post-mortems.
- Develop CI/CD pipelines with automated safety checks and rollback capabilities.
- Collaborate with product teams to advise on reliability and operational architecture.
View Full Description & ApplyYou'll be redirected to the employer's site