Senior MLOps Platform Engineer

New
J
JobgetherAI Infrastructure
Remote position within the United StatesFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
5+ years
Required Skills
AWSPythonJavaKubernetesData engineeringGoRustCI/CDDistributed Systems

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or related technical field.
  • 5+ years of experience building and operating production-grade software infrastructure.
  • Deep expertise with Kubernetes including Helm, operators, and container runtimes.
  • Hands-on experience with AWS services including EKS, SageMaker, S3, IAM, and CloudWatch.
  • Strong Python skills plus proficiency in Go, Rust, or Java.
  • Experience with CI/CD and GitOps tooling such as Argo CD, Flux, or GitLab.
  • Understanding of distributed systems, consensus, fault tolerance, and load balancing.
  • Experience with data engineering technologies like Airflow, Kafka, Spark, or Flink.
  • Familiarity with observability platforms and defining SLIs/SLOs.
  • Ability to collaborate with research and product teams on scalable AI services.

Responsibilities

  • Design, implement, and operate a unified MLOps platform spanning on-premise Kubernetes clusters and AWS.
  • Develop reusable GitLab CI and CI/CD pipelines for model packaging, containerization, and deployment.
  • Build and maintain observability and alerting capabilities using Prometheus, Grafana, OpenTelemetry, and CloudWatch.
  • Create self-service CLIs, SDKs, and dashboards for model registration and endpoint management.
  • Architect and maintain data pipelines for model artifacts and logs across S3 and on-premise storage.
  • Optimize inference performance through GPU/CPU scaling, quantization, and batching strategies.
  • Partner with research and engineering teams to productionize AI models with high reproducibility and security.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now