Generative AI Operations Engineer (GenAI Ops)

New
E
EPAMAI Operations
Opportunity to work remotely within PolandFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Languages
En B2
Experience
3+ years
Required Skills
AWSDockerPythonKubernetesCI/CDTerraformMLOpsGenerative AI

Requirements

  • 3+ years in a DevOps, SRE, or MLOps role with a focus on cloud infrastructure (AWS, GCP, or Azure).
  • Proficiency in building and managing CI/CD pipelines and at least one scripting language (Python or Bash).
  • Familiarity with IaC tools (e.g., AWS CDK, CloudFormation, Terraform) and containerization/orchestration (Docker, Kubernetes).
  • Proven track record of deploying and operating LLM inference (e.g., vLLM, Triton, TGI, Ray Serve, KServe/Seldon).
  • Hands-on experience with LLM/app tracing and metrics (e.g., OpenTelemetry, Langfuse, Arize Phoenix, WhyLabs).
  • Experience operating retrieval pipelines including embedding generation, indexing strategies, and vector databases (Pinecone, Weaviate, Milvus, FAISS).
  • Experience running multi-agent workflows (LangGraph, CrewAI, AutoGen) including state management and auditing.
  • Experience implementing security guardrails such as secrets isolation, prompt-injection defenses, and PII redaction.
  • B2 proficiency in English.

Responsibilities

  • Design, implement, and maintain robust, automated CI/CD pipelines for training, evaluating, and deploying LLMs and AI agents.
  • Design, deploy, and manage sophisticated, multi-agent systems ensuring seamless Agent-to-Agent (A2A) communication.
  • Implement and manage secure, scalable integrations between AI agents and external tools/APIs using standards like Model Context Protocol (MCP).
  • Utilize IaC services or tools like Terraform to define and manage infrastructure for GenAI workloads.
  • Implement comprehensive monitoring and logging solutions to track model and agent performance, resource utilization, and system health.
  • Design and implement scalable architectures for model serving and inference to optimize performance and cost.
  • Enforce security best practices, guardrails, and compliance standards for GenAI infrastructure and data.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now