Staff Site Reliability Engineer
New
S
SentinelOneCybersecurity software
Remote in CZ/SKFull-TimeStaff
SalarySalary starting from 4500 EUR/month. Annual bonus based on company performance, paid in two installments.
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years of experience in Site Reliability Engineering, DevOps, or a related field in cloud native environments
- Required Skills
- PythonBashKubernetesGoGrafanaPrometheusDevOps
Requirements
- Have 5+ years of experience in Site Reliability Engineering, DevOps, or a related field in cloud native environments.
- Have experience troubleshooting complex issues under pressure.
- Be able to use and improve established runbooks.
- Have experience with Kubernetes and container orchestration.
- Have experience with observability stacks such as Prometheus, Grafana, ELK, or OpenTelemetry.
- Be proficient in Python or Go and Bash scripting to improve operational workflows and incident response.
- Be familiar with modern CI/CD pipelines and DevOps practices.
- Have excellent communication skills and demonstrated ability to mentor peers in reliability practices.
Responsibilities
- Participate in and help drive incident management for production issues, ensuring rapid recovery and root cause analysis.
- Lead postmortems where needed and conduct post-incident reviews.
- Improve and optimize the observability strategy.
- Collaborate with application engineering teams to design and implement monitoring solutions that improve alerting and reduce noise.
- Help define and refine SLOs, SLIs, and SLAs aligned with business objectives and customer expectations.
- Document incident findings and drive follow-up actions to prevent recurrence.
- Mentor peers and junior engineers in incident response, troubleshooting techniques, and reliability best practices.
View Full Description & ApplyYou'll be redirected to the employer's site