Datacenter Infrastructure Specialist
New
R
RunpodAI Infrastructure
Remote - EMEAFull-TimeMiddle
Salary€105,324.00 to €140,432.00
Apply NowOpens the employer's application page
Job Details
- Experience
- 3–5 years
- Required Skills
- DockerPythonBashGoLinux
Requirements
- 3–5 years of experience in infrastructure operations, systems reliability, or datacenter engineering.
- Strong proficiency in standard datacenter networking and performance troubleshooting.
- Hands-on experience with the NVIDIA software stack (driver installation, performance utilities).
- Solid Linux system administration skills and experience with containerization (Docker).
- Comfortable performing system-level troubleshooting and performance tuning at the kernel and hardware interface layers.
- Effective written and verbal communication skills for technical and leadership audiences.
- Familiarity with RDMA, InfiniBand, or RoCE is highly preferred.
- Proficiency in Python, Go, or Bash for infrastructure automation (preferred).
- Experience with observability tools like Grafana, Prometheus, or Datadog (preferred).
- Experience managing bare-metal High-Performance Computing (HPC) environments (preferred).
- Experience in a fast-paced startup environment (preferred).
Responsibilities
- Validate new hardware to ensure partner deployments meet specifications for distributed AI/ML workloads.
- Monitor fleet health and provide technical data to protect customer SLAs and address performance degradation.
- Operate with an AI-first mindset, leveraging LLMs and AI agents to automate network triage and generate dynamic runbooks.
- Coordinate technical incident communications and translate outages into actionable resolutions.
- Support the ongoing growth and performance of infrastructure partners.
View Full Description & ApplyYou'll be redirected to the employer's site