Founding AI Platform Engineer (MLOps / Backend)
New
J
JobgetherAI Platform Engineering
IndiaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- PythonCI/CDSaaSLLMMLOpsGenerative AI
Requirements
- Strong software engineering background with experience building, deploying, and operating production systems.
- Proven experience with backend services, cloud infrastructure, CI/CD, automated testing, observability, and engineering automation.
- Strong proficiency in Python and the ability to work effectively across backend services, infrastructure, tooling, and operational workflows.
- Good understanding of reliability, performance, maintainability, scalability, and infrastructure cost tradeoffs.
- Ability to collaborate effectively with ML and product teams and independently drive ambiguous technical work to completion.
- Strong ownership mentality, attention to detail, and a practical approach focused on simplifying and strengthening systems.
- Experience with MLOps workflows covering model training, evaluation, deployment, and monitoring.
- Experience serving machine-learning models or LLM-powered applications in production.
- Familiarity with experimentation platforms, event pipelines, analytics instrumentation, or feature delivery platforms.
- Experience with agent evaluation, prompt versioning, retrieval and search infrastructure, or vector-backed systems.
- Experience supporting customer-facing APIs or SaaS platform infrastructure.
Responsibilities
- Build and maintain infrastructure and tooling for training, evaluating, deploying, serving, and monitoring ML models and GenAI services.
- Develop and operate production backend services, APIs, and pipelines supporting recommendations, agent workflows, and customer-facing integrations.
- Improve CI/CD pipelines, automated testing, release processes, rollback strategies, and environment management.
- Establish comprehensive observability across application health, model behavior, agent quality, latency, costs, and operational failure modes.
- Build reproducibility and lifecycle management practices for models, prompts, datasets, configurations, and software releases.
- Support experimentation and measurement infrastructure that enables ML and product teams to evaluate changes reliably.
- Strengthen platform reliability, scalability, security, performance, and cost efficiency across the technology stack.
- Collaborate closely with ML, product, and engineering teams to move ambiguous initiatives from concept to completion.
View Full Description & ApplyYou'll be redirected to the employer's site