Senior Site Reliability Engineer
New
H
Hard Rock DigitalOnline Gaming
Poland (Remote)Full-TimeSenior
Salary28,000 - 31,000 PLN per month
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years in SRE, DevOps, or similar infrastructure roles; 3+ years managing production Kubernetes; 1+ years building or operating AI/LLM-powered tools
- Required Skills
- DockerPythonJavaKubernetesAzureGrafanaDevOpsTerraformAnsible
Requirements
- 5+ years in SRE, DevOps, or similar infrastructure roles.
- 3+ years hands-on experience managing production Kubernetes clusters.
- 1+ years of practical experience building or operating AI/LLM-powered tools or workflows.
- Advanced expertise with the Grafana observability stack and PromQL.
- Experience managing Java-based applications, including JVM tuning and optimization.
- Proficiency with Infrastructure as Code tools such as Terraform or Ansible.
- Experience with agentic orchestration frameworks like LangChain, LangGraph, or CrewAI.
- Strong scripting abilities in Python, Bash, or Go.
- Experience with GitOps workflows and CI/CD pipelines.
- Degree in Computer Science or equivalent professional experience.
Responsibilities
- Ensure the availability, reliability, and performance of high-traffic Java-based applications in a distributed environment.
- Deploy and manage the Grafana stack (Grafana, Prometheus, Loki, Mimir, Alloy) to deliver real-time monitoring and alerting.
- Design and build autonomous AI agents that automate operational tasks like alert triage, root cause analysis, and runbook execution.
- Develop tool-calling LLM agents that interact with infrastructure APIs like Kubernetes, Jira, and Slack.
- Measure and report on toil reduction metrics to quantify the impact of automation initiatives.
- Support incident response efforts and conduct post-mortems using AI tools to surface patterns and accelerate timelines.
View Full Description & ApplyYou'll be redirected to the employer's site