Senior AI Infrastructure & Platform Operations Engineer
New
M
MirantisCloud Infrastructure
Remote in the USFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 7+ years
- Required Skills
- KubernetesLinuxNetworking
Requirements
- 7+ years of experience in infrastructure operations, platform operations, site reliability engineering, or related technical roles.
- Expert-level Linux administration and troubleshooting skills.
- Strong experience operating Kubernetes in production environments.
- Strong networking expertise including experience diagnosing complex performance and connectivity issues.
- Experience supporting large-scale production infrastructure and distributed systems.
- Proven experience leading technical investigations and managing complex incidents.
- Strong understanding of observability, monitoring, and service reliability practices.
- Excellent troubleshooting and analytical skills across multiple infrastructure domains.
Responsibilities
- Lead the investigation and resolution of complex infrastructure, networking, and platform-related incidents.
- Act as a senior escalation point for operational teams during critical service-impacting events.
- Support large-scale NVIDIA GPU infrastructure and high-performance networking environments.
- Identify opportunities to automate repetitive operational activities and improve operational efficiency.
- Mentor and support AI Infrastructure & Platform Operations Engineers.
View Full Description & ApplyYou'll be redirected to the employer's site