DevOps Engineer (AI Inference)
G
GcoreAI infrastructure
Workplace type: remote; Locations: Kraków, Warszawa, Gdańsk, Łódź, Poznań, Country code: PLFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Languages
- En B2
- Required Skills
- PythonBashKubernetesGoGrafanaLinuxTerraformAnsibleNetworking
Requirements
- Strong understanding of Kubernetes architecture, including CNI, CSI, operators, ingress/gateway, and control plane components.
- Hands-on experience operating and troubleshooting production Kubernetes clusters.
- Strong Linux and networking troubleshooting skills, including DNS, routing, firewalling, TLS, MTU, connectivity, and performance issues.
- Ability to develop automation and operational tooling using Python, Go, or Bash.
- Experience with Terraform, Ansible, or similar IaC/configuration management tools.
- Experience with VictoriaMetrics/Grafana or similar monitoring, alerting, and troubleshooting tools.
- Strong experience with Git-based workflows and CI/CD pipelines.
- Preferred: familiarity with Cluster API or similar Kubernetes cluster lifecycle management technologies.
- Preferred: hands-on operation or administration of Slurm clusters.
- Preferred: knowledge of Argo CD, GitOps workflows, Helm, or Helmfile.
- Preferred: exposure to bare metal, GPU, HPC, or other high-performance computing environments.
Responsibilities
- Design, develop, and maintain infrastructure for on-premises AI inference workloads.
- Implement GPU scheduling, model deployment pipelines, and data access patterns.
- Build monitoring and observability tools for AI inference platforms.
- Create dashboards, alerts, and runbooks for model health and system performance.
- Collaborate with ML engineers and platform teams on AI workload system architecture.
- Integrate inference runtimes and test performance at scale.
View Full Description & ApplyYou'll be redirected to the employer's site