Senior Site Reliability Engineer
New
P
PragmatikeCloud Computing
Fully remote EU timezone (CET ±2h). Eligible locations: Latvia, Spain, Albania, Bosnia & Herzegovina, Poland, Portugal, Italy., CET ±2hFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Languages
- Fluent English
- Required Skills
- PythonBashKubernetesGrafanaPrometheusLinuxAnsible
Requirements
- Expert-level, hands-on experience operating Kubernetes in production.
- Strong network engineering skills including VLANs, L2/L3 routing, and VPNs.
- Strong proficiency in Linux systems administration (Debian/Ubuntu).
- Experience building and maintaining automation workflows with Ansible, Bash, or Python.
- Proven experience with observability tools like Prometheus, Grafana, ELK, or Graylog.
- Background with virtualization technologies such as OpenStack, Proxmox, or VMware.
- Experience with bare-metal provisioning and MAAS.
- Solid understanding of distributed systems and container orchestration.
- Ability to develop SOPs and operational procedures from scratch.
- Experience managing incident response and on-call rotations.
Responsibilities
- Operate and maintain Linux-based infrastructure (Debian/Ubuntu).
- Deploy, manage, and scale Kubernetes clusters across bare-metal, virtualized, and on-prem environments.
- Oversee full cluster lifecycle including upgrades, networking, and security hardening.
- Implement automation for provisioning and operations using Ansible, Bash/Python, and GitOps workflows.
- Design and maintain network architecture including VLANs, L2/L3 routing, and VPNs.
- Deploy and maintain observability stacks like Prometheus, Grafana, Loki, ELK, or Graylog.
- Lead incident response and escalation activities across the platform.
- Define and implement SLOs/SLIs across infrastructure and software services.
View Full Description & ApplyYou'll be redirected to the employer's site