Senior Site Reliability Engineer / Kubernetes

New
P
PragmatikeCloud Computing
Location: Estonia; Secondary Locations: Latvia, Spain, Albania, Bosnia & Herzegovina, Poland, Portugal, Italy, Türkiye, Romania. Fully remote EU timezone (CET ±2h), CET ±2hFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Languages
Fluent English is mandatory
Required Skills
PythonBashKubernetesGrafanaPrometheusLinuxAnsible

Requirements

  • Expert-level, hands-on experience operating Kubernetes in production environments.
  • Strong network engineering skills, including VLANs, L2/L3 routing, VPNs, and multi-site connectivity.
  • Strong Linux systems administration proficiency with Debian and Ubuntu.
  • Ability to design complex network architectures and understanding of networking fundamentals.
  • Experience building and maintaining automation workflows using Ansible, Bash/Python, and Git-based tools.
  • Experience with observability stacks such as Prometheus, Grafana, ELK, Loki, or Graylog.
  • Background with OpenStack, Proxmox, or VMware virtualization technologies.
  • Experience with bare-metal provisioning and MAAS (Metal as a Service).
  • Strong understanding of distributed systems and container orchestration.
  • Experience with incident response, escalation procedures, and on-call rotations.
  • Ability to develop SOPs and operational procedures from scratch.
  • Fluent English.

Responsibilities

  • Operate and maintain Linux-based infrastructure running Debian and Ubuntu.
  • Deploy, manage, and scale Kubernetes clusters across bare-metal, virtualized, and on-prem environments.
  • Oversee cluster upgrades, node pools, networking, storage, and security hardening.
  • Implement provisioning and operations automation using Ansible, Bash/Python, and GitOps workflows.
  • Design and maintain networking architecture, including VLANs, L2/L3 routing, VPNs, and multi-site connectivity.
  • Build deployment workflows using PXE boot, Preseed, and cloud-init.
  • Deploy and maintain observability stacks including Prometheus/Grafana, Loki, ELK, and Graylog.
  • Lead incident response and escalation; improve availability, latency, alerting, and monitoring.
  • Define SLOs and SLIs, establish on-call schedules, and develop standard operating procedures.
  • Manage virtualization and orchestration layers, coordinate physical maintenance, and collaborate with development teams and cross-functional stakeholders.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now