Senior Site Reliability Engineer II
New
S
SmartRecruiters Inc.AI Hiring Platform
You may be located anywhere in Poland and work remotely or out of our Cracow office.ContractSenior
Salary4940 - 6916 GBP per month gross permanent currencySource=conversion; 6690 - 9366 USD per month gross permanent currencySource=conversion; 25000 - 35000 PLN per month gross permanent currencySource=original; 5419 - 7587 CHF per month gross permanent currencySource=conversion; 5770 - 8078 EUR per month gross permanent currencySource=conversion
Apply NowOpens the employer's application page
Job Details
- Languages
- English (C1 level)
- Experience
- 7+ years of professional experience
- Required Skills
- AWSPythonJavaKubernetesGoLinuxMicroservicesDistributed Systems
Requirements
- 7+ years of professional experience.
- Strong working knowledge of SRE and observability industry standards (SLIs/SLOs, error budgets, incident management, on-call).
- Software development experience in JVM stack.
- Experience with AWS, Kubernetes, and IaC tools for running and troubleshooting distributed applications.
- Deep knowledge of Linux operating system, including system hardening and performance troubleshooting.
- Very good scripting skills in Bash, Golang, or Python.
- Strong English verbal and written communication skills.
- Ability to coach and influence other engineers.
- Understanding of networking standards (TCP/IP, DNS, VPN, load balancing) is a plus.
- Experience managing and troubleshooting SQL and NoSQL database systems is a plus.
Responsibilities
- Improve, automate, and grow observability and reliability tooling including metrics, logs, traces, and alerting.
- Mentor engineering team members and advocate for modern SRE practices.
- Partner with product engineers (Java, Node.js, Python) to design, instrument, and operate services, managing SLIs/SLOs and error budgets.
- Create reusable building blocks such as dashboards, alerts, libraries, and IaC modules for company-wide use.
- Respond to production incidents and threats, lead remediation, and drive follow-up improvements.
- Document standards, best practices, and policies for monitoring, alerting, incident response, and reliability.
- Lead reliability-related initiatives across the organization.
View Full Description & ApplyYou'll be redirected to the employer's site