DevOps Engineer (AI Inference)

G
GcoreAI infrastructure
Workplace type: remote; Locations: Kraków, Warszawa, Gdańsk, Łódź, Poznań, Country code: PLFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Languages
En B2
Required Skills
PythonBashKubernetesGoGrafanaLinuxTerraformAnsibleNetworking

Requirements

  • Strong understanding of Kubernetes architecture, including CNI, CSI, operators, ingress/gateway, and control plane components.
  • Hands-on experience operating and troubleshooting production Kubernetes clusters.
  • Strong Linux and networking troubleshooting skills, including DNS, routing, firewalling, TLS, MTU, connectivity, and performance issues.
  • Ability to develop automation and operational tooling using Python, Go, or Bash.
  • Experience with Terraform, Ansible, or similar IaC/configuration management tools.
  • Experience with VictoriaMetrics/Grafana or similar monitoring, alerting, and troubleshooting tools.
  • Strong experience with Git-based workflows and CI/CD pipelines.
  • Preferred: familiarity with Cluster API or similar Kubernetes cluster lifecycle management technologies.
  • Preferred: hands-on operation or administration of Slurm clusters.
  • Preferred: knowledge of Argo CD, GitOps workflows, Helm, or Helmfile.
  • Preferred: exposure to bare metal, GPU, HPC, or other high-performance computing environments.

Responsibilities

  • Design, develop, and maintain infrastructure for on-premises AI inference workloads.
  • Implement GPU scheduling, model deployment pipelines, and data access patterns.
  • Build monitoring and observability tools for AI inference platforms.
  • Create dashboards, alerts, and runbooks for model health and system performance.
  • Collaborate with ML engineers and platform teams on AI workload system architecture.
  • Integrate inference runtimes and test performance at scale.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now