Senior Site Reliability Engineer (AI Hardware & Infrastructure)
New
L
Link GroupAI Infrastructure
Warszawa, -, Gdańsk, -, Kraków, -, Wrocław, -, Poznań, -, Lublin, -, Łódź, -, Białystok, -, Olsztyn, -, Szczecin, -, Warszawa, Country code: PLFull-TimeSenior
Salary20,000 - 27,000 PLN per month
Apply NowOpens the employer's application page
Job Details
- Required Skills
- PythonArtificial IntelligenceGrafanaPrometheusNetworking
Requirements
- Deep background in Site Reliability or Production Engineering.
- Strong Computer Science foundation.
- Proven experience managing large-scale, mission-critical infrastructure.
- Exceptional Python programming skills with experience building scalable operational tools.
- Practical understanding of advanced network topologies, high-bandwidth routing/switching, BGP, and dual-stack IPv4/IPv6.
- Expert-level hands-on experience with Prometheus, Grafana, OpenTelemetry, and Loki.
- Experience designing service rollout strategies and defining alerting thresholds.
- Experience creating technical runbooks and leading incident response teams.
- Ability to partner effectively with external data center vendors and on-site technicians.
Responsibilities
- Write Python automation to manage the lifecycle of server fleet from provisioning to decommissioning.
- Integrate operational systems including JIRA, Siebel, and PagerDuty via APIs to automate incident resolution.
- Design and implement observability stacks using Prometheus, Grafana, OpenTelemetry, and Loki with custom telemetry pipelines.
- Lead incident response for high-severity outages and conduct blameless post-mortems.
- Participate in a 24/7 on-call rotation.
- Utilize AI utilities and LLM-assisted development to enhance technical execution and performance evaluation.
- Coordinate with data center vendors and on-site technicians to ensure infrastructure uptime.
View Full Description & ApplyYou'll be redirected to the employer's site