Senior Site Reliability Engineer (MAAS)

New
P
PragmatikeCloud computing, GPU infrastructure
Location: Armenia Secondary Locations: Bulgaria, Spain, Albania, Ukraine, Greece, Croatia, Serbia, Poland, Moldova, Portugal, Italy, Türkiye, Montenegro, Romania Workplace: Remote, EU timezone (CET ±2h)Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Languages
Fluent English required
Experience
5+ years of hands-on SRE, Infrastructure, Systems, or Platform Engineering experience.
Required Skills
PythonBashKubernetesGrafanaPrometheusTerraformAnsible

Requirements

  • Have 5+ years of hands-on SRE, infrastructure, systems, or platform engineering experience.
  • Bring expert-level Linux administration experience, particularly with Debian/Ubuntu.
  • Have strong production experience with MAAS and bare-metal provisioning.
  • Have expert-level, hands-on experience operating Kubernetes in production, including cluster lifecycle, networking, storage, upgrades, and troubleshooting.
  • Have network engineering skills across VLANs, L2/L3 routing, bonding, VPNs, firewalls, and DNS.
  • Have strong automation skills with Ansible and Bash and/or Python.
  • Have experience with Terraform/OpenTofu and Git-based infrastructure workflows.
  • Have production experience with Prometheus/Grafana or comparable observability platforms.
  • Have experience with incident response, on-call operations, monitoring, alerting, and reliability practices.
  • Have experience with virtualization technologies such as Proxmox, KVM/libvirt, OpenStack, or VMware.
  • Have experience with bare-metal hardware, BMCs, IPMI/Redfish, storage, and hardware troubleshooting.
  • Understand distributed systems, container orchestration, infrastructure reliability, and infrastructure security.

Responsibilities

  • Operate and maintain large-scale Linux infrastructure across Debian/Ubuntu-based bare-metal and virtualized environments.
  • Own MAAS-based bare-metal provisioning, including controllers, PXE, commissioning, cloud-init, node lifecycle, and API/CLI automation.
  • Operate and maintain production Kubernetes clusters, including upgrades, networking, storage, security hardening, and troubleshooting.
  • Design and maintain multi-site networking across VLANs, L2/L3 routing, VPNs, firewalls, and DNS.
  • Automate infrastructure provisioning and operations using Ansible, Bash/Python, OpenTofu/Terraform, and Git-based workflows.
  • Operate observability platforms and define SLIs, SLOs, alerting, and reliability practices.
  • Lead infrastructure incident response, troubleshooting, escalation, and post-incident improvements.
  • Manage virtualization platforms and work with hardware-layer systems, including BMCs, storage, and GPU infrastructure.
  • Build internal infrastructure tooling and own site onboarding, maintenance, decommissioning, drift detection, and operational runbooks.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now