Founding AI Platform Engineer (MLOps / Backend)

New
J
JobgetherAI Platform Engineering
IndiaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Required Skills
PythonCI/CDSaaSLLMMLOpsGenerative AI

Requirements

  • Strong software engineering background with experience building, deploying, and operating production systems.
  • Proven experience with backend services, cloud infrastructure, CI/CD, automated testing, observability, and engineering automation.
  • Strong proficiency in Python and the ability to work effectively across backend services, infrastructure, tooling, and operational workflows.
  • Good understanding of reliability, performance, maintainability, scalability, and infrastructure cost tradeoffs.
  • Ability to collaborate effectively with ML and product teams and independently drive ambiguous technical work to completion.
  • Strong ownership mentality, attention to detail, and a practical approach focused on simplifying and strengthening systems.
  • Experience with MLOps workflows covering model training, evaluation, deployment, and monitoring.
  • Experience serving machine-learning models or LLM-powered applications in production.
  • Familiarity with experimentation platforms, event pipelines, analytics instrumentation, or feature delivery platforms.
  • Experience with agent evaluation, prompt versioning, retrieval and search infrastructure, or vector-backed systems.
  • Experience supporting customer-facing APIs or SaaS platform infrastructure.

Responsibilities

  • Build and maintain infrastructure and tooling for training, evaluating, deploying, serving, and monitoring ML models and GenAI services.
  • Develop and operate production backend services, APIs, and pipelines supporting recommendations, agent workflows, and customer-facing integrations.
  • Improve CI/CD pipelines, automated testing, release processes, rollback strategies, and environment management.
  • Establish comprehensive observability across application health, model behavior, agent quality, latency, costs, and operational failure modes.
  • Build reproducibility and lifecycle management practices for models, prompts, datasets, configurations, and software releases.
  • Support experimentation and measurement infrastructure that enables ML and product teams to evaluate changes reliably.
  • Strengthen platform reliability, scalability, security, performance, and cost efficiency across the technology stack.
  • Collaborate closely with ML, product, and engineering teams to move ambiguous initiatives from concept to completion.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now