Senior Site Reliability Engineer / Kubernetes
New
P
PragmatikeCloud Computing
Location: Estonia; Secondary Locations: Latvia, Spain, Albania, Bosnia & Herzegovina, Poland, Portugal, Italy, Türkiye, Romania. Fully remote EU timezone (CET ±2h), CET ±2hFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Languages
- Fluent English is mandatory
- Required Skills
- PythonBashKubernetesGrafanaPrometheusLinuxAnsible
Requirements
- Expert-level, hands-on experience operating Kubernetes in production environments.
- Strong network engineering skills, including VLANs, L2/L3 routing, VPNs, and multi-site connectivity.
- Strong Linux systems administration proficiency with Debian and Ubuntu.
- Ability to design complex network architectures and understanding of networking fundamentals.
- Experience building and maintaining automation workflows using Ansible, Bash/Python, and Git-based tools.
- Experience with observability stacks such as Prometheus, Grafana, ELK, Loki, or Graylog.
- Background with OpenStack, Proxmox, or VMware virtualization technologies.
- Experience with bare-metal provisioning and MAAS (Metal as a Service).
- Strong understanding of distributed systems and container orchestration.
- Experience with incident response, escalation procedures, and on-call rotations.
- Ability to develop SOPs and operational procedures from scratch.
- Fluent English.
Responsibilities
- Operate and maintain Linux-based infrastructure running Debian and Ubuntu.
- Deploy, manage, and scale Kubernetes clusters across bare-metal, virtualized, and on-prem environments.
- Oversee cluster upgrades, node pools, networking, storage, and security hardening.
- Implement provisioning and operations automation using Ansible, Bash/Python, and GitOps workflows.
- Design and maintain networking architecture, including VLANs, L2/L3 routing, VPNs, and multi-site connectivity.
- Build deployment workflows using PXE boot, Preseed, and cloud-init.
- Deploy and maintain observability stacks including Prometheus/Grafana, Loki, ELK, and Graylog.
- Lead incident response and escalation; improve availability, latency, alerting, and monitoring.
- Define SLOs and SLIs, establish on-call schedules, and develop standard operating procedures.
- Manage virtualization and orchestration layers, coordinate physical maintenance, and collaborate with development teams and cross-functional stakeholders.
View Full Description & ApplyYou'll be redirected to the employer's site