Senior Infrastructure Engineer - AI/ML Platform

New
O
OpenTeamsAI/ML Infrastructure
U.S - RemoteFull-TimeSenior
Salary$145,000–$250,000 USD, dependent on experience level and location
Apply NowOpens the employer's application page

Job Details

Experience
6+ years
Required Skills
AWSPythonKubernetesAzureGoCI/CDDevOpsTerraform

Requirements

  • U.S. citizenship required.
  • Ability to obtain and maintain a Secret security clearance.
  • 6+ years of hands-on infrastructure, platform, DevOps, or site reliability engineering experience.
  • Strong understanding of infrastructure engineering principles (scalability, reliability, observability, security, automation).
  • Production experience with Kubernetes (workload scheduling, resource management, multi-tenant environments).
  • Experience with automated security tooling (container scanning, SAST/DAST, artifact signing, policy enforcement).
  • Proficiency with infrastructure-as-code tools (Terraform, OpenTofu, Pulumi, or similar).
  • Experience with at least one major cloud platform (AWS, Azure, or Google Cloud).
  • Experience implementing monitoring and observability (OpenTelemetry, Prometheus, Grafana, or similar).
  • Strong programming or automation skills using Python, Go, or a comparable language.
  • Experience with CI/CD practices and GitOps workflows.
  • Experience creating operational documentation and runbooks.
  • Experience leading technical initiatives or mentoring other engineers.

Responsibilities

  • Build and operate the Kubernetes platform supporting AI test and evaluation frameworks.
  • Implement GPU scheduling, workload orchestration, resource management, and multi-tenant isolation.
  • Design infrastructure-as-code, GitOps workflows, and automated deployment pipelines.
  • Develop reusable and modular infrastructure components.
  • Contribute to Nebari and other open-source Kubernetes and MLOps projects.
  • Own platform reliability, including capacity planning, upgrade strategies, and failure-mode analysis.
  • Design and implement observability, monitoring, logging, tracing, and alerting for large-scale AI/ML workloads.
  • Develop operational runbooks and documentation for deployment and troubleshooting.
  • Deploy, configure, and harden infrastructure within secure, restricted, or disconnected Government environments.
  • Support security authorization and compliance activities through documentation and hardened configurations.
View Full Description & ApplyYou'll be redirected to the employer's site
$145,000–$250,000 USD, dependent on experience level and location
Apply Now