Senior Site Reliability Engineer

New
P
PragmatikeCloud Computing
Fully remote EU timezone (CET ±2h). Eligible locations: Latvia, Spain, Albania, Bosnia & Herzegovina, Poland, Portugal, Italy., CET ±2hFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Languages
Fluent English
Required Skills
PythonBashKubernetesGrafanaPrometheusLinuxAnsible

Requirements

  • Expert-level, hands-on experience operating Kubernetes in production.
  • Strong network engineering skills including VLANs, L2/L3 routing, and VPNs.
  • Strong proficiency in Linux systems administration (Debian/Ubuntu).
  • Experience building and maintaining automation workflows with Ansible, Bash, or Python.
  • Proven experience with observability tools like Prometheus, Grafana, ELK, or Graylog.
  • Background with virtualization technologies such as OpenStack, Proxmox, or VMware.
  • Experience with bare-metal provisioning and MAAS.
  • Solid understanding of distributed systems and container orchestration.
  • Ability to develop SOPs and operational procedures from scratch.
  • Experience managing incident response and on-call rotations.

Responsibilities

  • Operate and maintain Linux-based infrastructure (Debian/Ubuntu).
  • Deploy, manage, and scale Kubernetes clusters across bare-metal, virtualized, and on-prem environments.
  • Oversee full cluster lifecycle including upgrades, networking, and security hardening.
  • Implement automation for provisioning and operations using Ansible, Bash/Python, and GitOps workflows.
  • Design and maintain network architecture including VLANs, L2/L3 routing, and VPNs.
  • Deploy and maintain observability stacks like Prometheus, Grafana, Loki, ELK, or Graylog.
  • Lead incident response and escalation activities across the platform.
  • Define and implement SLOs/SLIs across infrastructure and software services.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now