Senior Infrastructure Engineer - AI/ML Platform
New
O
OpenTeamsAI/ML Infrastructure
U.S - RemoteFull-TimeSenior
Salary$145,000–$250,000 USD, dependent on experience level and location
Apply NowOpens the employer's application page
Job Details
- Experience
- 6+ years
- Required Skills
- AWSPythonKubernetesAzureGoCI/CDDevOpsTerraform
Requirements
- U.S. citizenship required.
- Ability to obtain and maintain a Secret security clearance.
- 6+ years of hands-on infrastructure, platform, DevOps, or site reliability engineering experience.
- Strong understanding of infrastructure engineering principles (scalability, reliability, observability, security, automation).
- Production experience with Kubernetes (workload scheduling, resource management, multi-tenant environments).
- Experience with automated security tooling (container scanning, SAST/DAST, artifact signing, policy enforcement).
- Proficiency with infrastructure-as-code tools (Terraform, OpenTofu, Pulumi, or similar).
- Experience with at least one major cloud platform (AWS, Azure, or Google Cloud).
- Experience implementing monitoring and observability (OpenTelemetry, Prometheus, Grafana, or similar).
- Strong programming or automation skills using Python, Go, or a comparable language.
- Experience with CI/CD practices and GitOps workflows.
- Experience creating operational documentation and runbooks.
- Experience leading technical initiatives or mentoring other engineers.
Responsibilities
- Build and operate the Kubernetes platform supporting AI test and evaluation frameworks.
- Implement GPU scheduling, workload orchestration, resource management, and multi-tenant isolation.
- Design infrastructure-as-code, GitOps workflows, and automated deployment pipelines.
- Develop reusable and modular infrastructure components.
- Contribute to Nebari and other open-source Kubernetes and MLOps projects.
- Own platform reliability, including capacity planning, upgrade strategies, and failure-mode analysis.
- Design and implement observability, monitoring, logging, tracing, and alerting for large-scale AI/ML workloads.
- Develop operational runbooks and documentation for deployment and troubleshooting.
- Deploy, configure, and harden infrastructure within secure, restricted, or disconnected Government environments.
- Support security authorization and compliance activities through documentation and hardened configurations.
View Full Description & ApplyYou'll be redirected to the employer's site