Senior AI Infrastructure & Platform Operations Engineer

New
M
MirantisCloud Infrastructure
Remote in the USFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
7+ years
Required Skills
KubernetesLinuxNetworking

Requirements

  • 7+ years of experience in infrastructure operations, platform operations, site reliability engineering, or related technical roles.
  • Expert-level Linux administration and troubleshooting skills.
  • Strong experience operating Kubernetes in production environments.
  • Strong networking expertise including experience diagnosing complex performance and connectivity issues.
  • Experience supporting large-scale production infrastructure and distributed systems.
  • Proven experience leading technical investigations and managing complex incidents.
  • Strong understanding of observability, monitoring, and service reliability practices.
  • Excellent troubleshooting and analytical skills across multiple infrastructure domains.

Responsibilities

  • Lead the investigation and resolution of complex infrastructure, networking, and platform-related incidents.
  • Act as a senior escalation point for operational teams during critical service-impacting events.
  • Support large-scale NVIDIA GPU infrastructure and high-performance networking environments.
  • Identify opportunities to automate repetitive operational activities and improve operational efficiency.
  • Mentor and support AI Infrastructure & Platform Operations Engineers.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now