Datacenter Infrastructure Specialist

New
R
RunpodAI Infrastructure
Remote - EMEAFull-TimeMiddle
Salary€105,324.00 to €140,432.00
Apply NowOpens the employer's application page

Job Details

Experience
3–5 years
Required Skills
DockerPythonBashGoLinux

Requirements

  • 3–5 years of experience in infrastructure operations, systems reliability, or datacenter engineering.
  • Strong proficiency in standard datacenter networking and performance troubleshooting.
  • Hands-on experience with the NVIDIA software stack (driver installation, performance utilities).
  • Solid Linux system administration skills and experience with containerization (Docker).
  • Comfortable performing system-level troubleshooting and performance tuning at the kernel and hardware interface layers.
  • Effective written and verbal communication skills for technical and leadership audiences.
  • Familiarity with RDMA, InfiniBand, or RoCE is highly preferred.
  • Proficiency in Python, Go, or Bash for infrastructure automation (preferred).
  • Experience with observability tools like Grafana, Prometheus, or Datadog (preferred).
  • Experience managing bare-metal High-Performance Computing (HPC) environments (preferred).
  • Experience in a fast-paced startup environment (preferred).

Responsibilities

  • Validate new hardware to ensure partner deployments meet specifications for distributed AI/ML workloads.
  • Monitor fleet health and provide technical data to protect customer SLAs and address performance degradation.
  • Operate with an AI-first mindset, leveraging LLMs and AI agents to automate network triage and generate dynamic runbooks.
  • Coordinate technical incident communications and translate outages into actionable resolutions.
  • Support the ongoing growth and performance of infrastructure partners.
View Full Description & ApplyYou'll be redirected to the employer's site
€105,324.00 to €140,432.00
Apply Now