Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure)
J
JobgetherGPU Cloud Infrastructure
GermanyFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- DockerPythonKubernetesLinuxTerraformAnsible
Requirements
- Strong hands-on experience deploying and maintaining datacenter infrastructure (GPU, HPC, AI, or private cloud).
- Proven experience with NVIDIA GPU servers, firmware, drivers, and PCIe topology.
- Strong Linux troubleshooting capabilities (OS, kernels, drivers, and hardware interfaces).
- Practical knowledge of networking (VLANs, VRFs, BGP, ECMP, OVS/OVN, and high-speed connectivity).
- Experience with GPU networking technologies like RoCE/RDMA, SR-IOV, and BlueField DPUs.
- Experience with virtualization and container platforms such as KVM/QEMU, Docker/containerd, and Kubernetes.
- Experience with distributed storage platforms and local NVMe management.
- Proficiency in automation tools such as Terraform, Ansible, Bash, or Python.
- Familiarity with infrastructure observability and telemetry tools (Prometheus, Grafana, Zabbix).
- Strong attention to detail in producing technical documentation and operational runbooks.
- Ability to assess physical datacenter requirements (power density, cooling, floor loading).
- Strong problem-solving and systems-thinking ability across hardware, network, and software layers.
Responsibilities
- Coordinate datacenter deployments, including physical rack layouts, power, and cooling requirements.
- Commission GPU servers, storage nodes, and platform infrastructure, ensuring validation of BIOS, firmware, and PCIe topology.
- Validate high-speed fiber/copper cabling and network connectivity, including switch configurations and BGP routing.
- Support RoCE/RDMA fabric validation for distributed AI workloads.
- Install and configure Linux/Ubuntu environments, NVIDIA drivers, CUDA, and container runtimes.
- Perform day-to-day operations, maintenance, and incident resolution for production infrastructure.
- Maintain accurate as-built documentation, runbooks, and deployment readiness reports.
- Partner with cross-functional engineering teams to drive infrastructure improvements and operational excellence.
View Full Description & ApplyYou'll be redirected to the employer's site