Senior Site Reliability Engineer (AI Hardware & Infrastructure)

New
L
Link GroupAI Infrastructure
Warszawa, -, Gdańsk, -, Kraków, -, Wrocław, -, Poznań, -, Lublin, -, Łódź, -, Białystok, -, Olsztyn, -, Szczecin, -, Warszawa, Country code: PLFull-TimeSenior
Salary20,000 - 27,000 PLN per month
Apply NowOpens the employer's application page

Job Details

Required Skills
PythonArtificial IntelligenceGrafanaPrometheusNetworking

Requirements

  • Deep background in Site Reliability or Production Engineering.
  • Strong Computer Science foundation.
  • Proven experience managing large-scale, mission-critical infrastructure.
  • Exceptional Python programming skills with experience building scalable operational tools.
  • Practical understanding of advanced network topologies, high-bandwidth routing/switching, BGP, and dual-stack IPv4/IPv6.
  • Expert-level hands-on experience with Prometheus, Grafana, OpenTelemetry, and Loki.
  • Experience designing service rollout strategies and defining alerting thresholds.
  • Experience creating technical runbooks and leading incident response teams.
  • Ability to partner effectively with external data center vendors and on-site technicians.

Responsibilities

  • Write Python automation to manage the lifecycle of server fleet from provisioning to decommissioning.
  • Integrate operational systems including JIRA, Siebel, and PagerDuty via APIs to automate incident resolution.
  • Design and implement observability stacks using Prometheus, Grafana, OpenTelemetry, and Loki with custom telemetry pipelines.
  • Lead incident response for high-severity outages and conduct blameless post-mortems.
  • Participate in a 24/7 on-call rotation.
  • Utilize AI utilities and LLM-assisted development to enhance technical execution and performance evaluation.
  • Coordinate with data center vendors and on-site technicians to ensure infrastructure uptime.
View Full Description & ApplyYou'll be redirected to the employer's site
20,000 - 27,000 PLN per month
Apply Now