Senior Site Reliability Engineer
New
P
PlaysonIGaming
Warszawa, PolandContractSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- AWSDockerNode.jsPythonGitKubernetesGoGrafanaPrometheusTerraform
Requirements
- Strong hands-on experience with Kubernetes (deployment, scaling, troubleshooting) in high-load environments
- Experience with GitOps tools such as FluxCD or ArgoCD
- Proven experience in incident response, root cause analysis, and postmortems in production systems
- Solid experience with AWS, Terraform, Docker, and CI/CD pipelines
- Experience with monitoring and observability tools such as Datadog, Prometheus, Grafana, and logging stacks like ELK or CloudWatch
- Strong understanding of networking concepts and protocols
- Proficiency in at least one scripting language (e.g. Python, Go, Node.js)
- Experience working with version control systems (Git)
- Familiarity with incident management tools like PagerDuty, Opsgenie, or similar
- Ability to operate effectively in a fast-paced, high-pressure environment with strong ownership and accountability
- Proactive, resilient mindset with a focus on continuous improvement and system stability
Responsibilities
- Own system reliability by actively monitoring platform health, managing alerts, and responding to incidents in real time
- Participate in 24/7 on-call rotations, taking full ownership of production stability in a high-traffic (5–7k RPS) environment
- Investigate incidents, perform root cause analysis, and implement long-term fixes to prevent recurrence
- Build and continuously improve monitoring, alerting, and observability across the Kubernetes (EKS) ecosystem
- Deploy, manage, and optimise infrastructure using Terraform, Helm, and GitOps tools (Flux/ArgoCD)
- Drive automation and proactively improve system resilience, reducing manual intervention and recurring issues
- Maintain and evolve CI/CD pipelines and infrastructure-as-code practices
- Collaborate closely with engineering teams to support deployments and minimise user impact in a live environment
- Introduce and integrate new tools and technologies to enhance scalability, reliability, and performance
- Handle environment-specific requests and ensure smooth day-to-day platform operations under constant load
View Full Description & ApplyYou'll be redirected to the employer's site