Senior Site Reliability Engineer (MAAS)
New
P
PragmatikeCloud computing, GPU infrastructure
Location: Armenia
Secondary Locations: Bulgaria, Spain, Albania, Ukraine, Greece, Croatia, Serbia, Poland, Moldova, Portugal, Italy, Türkiye, Montenegro, Romania
Workplace: Remote, EU timezone (CET ±2h)Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Languages
- Fluent English required
- Experience
- 5+ years of hands-on SRE, Infrastructure, Systems, or Platform Engineering experience.
- Required Skills
- PythonBashKubernetesGrafanaPrometheusTerraformAnsible
Requirements
- Have 5+ years of hands-on SRE, infrastructure, systems, or platform engineering experience.
- Bring expert-level Linux administration experience, particularly with Debian/Ubuntu.
- Have strong production experience with MAAS and bare-metal provisioning.
- Have expert-level, hands-on experience operating Kubernetes in production, including cluster lifecycle, networking, storage, upgrades, and troubleshooting.
- Have network engineering skills across VLANs, L2/L3 routing, bonding, VPNs, firewalls, and DNS.
- Have strong automation skills with Ansible and Bash and/or Python.
- Have experience with Terraform/OpenTofu and Git-based infrastructure workflows.
- Have production experience with Prometheus/Grafana or comparable observability platforms.
- Have experience with incident response, on-call operations, monitoring, alerting, and reliability practices.
- Have experience with virtualization technologies such as Proxmox, KVM/libvirt, OpenStack, or VMware.
- Have experience with bare-metal hardware, BMCs, IPMI/Redfish, storage, and hardware troubleshooting.
- Understand distributed systems, container orchestration, infrastructure reliability, and infrastructure security.
Responsibilities
- Operate and maintain large-scale Linux infrastructure across Debian/Ubuntu-based bare-metal and virtualized environments.
- Own MAAS-based bare-metal provisioning, including controllers, PXE, commissioning, cloud-init, node lifecycle, and API/CLI automation.
- Operate and maintain production Kubernetes clusters, including upgrades, networking, storage, security hardening, and troubleshooting.
- Design and maintain multi-site networking across VLANs, L2/L3 routing, VPNs, firewalls, and DNS.
- Automate infrastructure provisioning and operations using Ansible, Bash/Python, OpenTofu/Terraform, and Git-based workflows.
- Operate observability platforms and define SLIs, SLOs, alerting, and reliability practices.
- Lead infrastructure incident response, troubleshooting, escalation, and post-incident improvements.
- Manage virtualization platforms and work with hardware-layer systems, including BMCs, storage, and GPU infrastructure.
- Build internal infrastructure tooling and own site onboarding, maintenance, decommissioning, drift detection, and operational runbooks.
View Full Description & ApplyYou'll be redirected to the employer's site