Senior MLOps Platform Engineer
New
J
JobgetherAI Infrastructure
Remote position within the United StatesFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- AWSPythonJavaKubernetesData engineeringGoRustCI/CDDistributed Systems
Requirements
- Bachelor’s degree in Computer Science, Engineering, or related technical field.
- 5+ years of experience building and operating production-grade software infrastructure.
- Deep expertise with Kubernetes including Helm, operators, and container runtimes.
- Hands-on experience with AWS services including EKS, SageMaker, S3, IAM, and CloudWatch.
- Strong Python skills plus proficiency in Go, Rust, or Java.
- Experience with CI/CD and GitOps tooling such as Argo CD, Flux, or GitLab.
- Understanding of distributed systems, consensus, fault tolerance, and load balancing.
- Experience with data engineering technologies like Airflow, Kafka, Spark, or Flink.
- Familiarity with observability platforms and defining SLIs/SLOs.
- Ability to collaborate with research and product teams on scalable AI services.
Responsibilities
- Design, implement, and operate a unified MLOps platform spanning on-premise Kubernetes clusters and AWS.
- Develop reusable GitLab CI and CI/CD pipelines for model packaging, containerization, and deployment.
- Build and maintain observability and alerting capabilities using Prometheus, Grafana, OpenTelemetry, and CloudWatch.
- Create self-service CLIs, SDKs, and dashboards for model registration and endpoint management.
- Architect and maintain data pipelines for model artifacts and logs across S3 and on-premise storage.
- Optimize inference performance through GPU/CPU scaling, quantization, and batching strategies.
- Partner with research and engineering teams to productionize AI models with high reproducibility and security.
View Full Description & ApplyYou'll be redirected to the employer's site