Senior AI Infrastructure & Platform Operations Engineer
New
M
MirantisAI Infrastructure
Remote in the EUFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 7+ years
- Required Skills
- KubernetesLinuxDistributed Systems
Requirements
- 7+ years of experience in infrastructure operations, platform operations, site reliability engineering, network operations, cloud operations, datacenter operations, or related technical roles.
- Expert-level Linux administration and troubleshooting skills.
- Strong networking expertise, including experience diagnosing complex performance, connectivity, and reliability issues.
- Strong experience operating Kubernetes in production environments.
- Experience supporting large-scale production infrastructure and distributed systems.
- Proven experience leading technical investigations and managing complex incidents.
- Experience performing root cause analysis and driving long-term operational improvements.
- Strong understanding of observability, monitoring, and service reliability practices.
- Excellent troubleshooting and analytical skills across multiple infrastructure domains.
- Strong communication, collaboration, and stakeholder management skills.
Responsibilities
- Lead the investigation and resolution of complex infrastructure, networking, and platform-related incidents.
- Act as a senior escalation point for operational teams during critical service-impacting events.
- Support large-scale NVIDIA GPU infrastructure and high-performance networking environments.
- Troubleshoot complex Linux, Kubernetes, networking, storage, and hardware-related issues.
- Analyze platform performance, capacity, stability, and reliability trends to proactively identify risks.
- Lead root cause analysis activities and drive long-term corrective actions.
- Provide technical leadership for Kubernetes platform operations and supporting infrastructure services.
- Drive improvements in platform reliability, observability, monitoring, and operational processes.
- Identify opportunities to automate repetitive operational activities and improve operational efficiency.
- Mentor and support AI Infrastructure & Platform Operations Engineers.
View Full Description & ApplyYou'll be redirected to the employer's site