Senior Data Engineer (AI/ML)

New
J
JobgetherSecurity & IT
100% remote position across India., Flexible collaboration across international teams and time zones.Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
5+ years of experience in data engineering, software engineering, distributed systems, or a related field.
Required Skills
PythonSQLJavaSnowflakeAirflowSparkScalaDatabricksGenerative AI

Requirements

  • 5+ years of experience in data engineering, software engineering, distributed systems, or a related field.
  • Strong programming skills in Python and/or Scala/Java, combined with advanced SQL capabilities.
  • Hands-on experience with Databricks, Snowflake, Apache Spark, Delta Lake, and Airflow.
  • Strong experience working with cloud-based data platforms and scalable data architectures.
  • Practical experience building applications using LLMs or Generative AI.
  • Strong understanding of RAG architectures, embeddings, vector databases, semantic search, and retrieval systems.
  • Experience with real-time and streaming architectures using technologies such as Kafka or Spark Structured Streaming.
  • Experience with LLM/AI evaluation frameworks, automated evaluations, experimentation, and quality metrics.
  • Familiarity with AI observability and tracing, including production monitoring.
  • Experience with LangGraph, LangChain, LlamaIndex, or similar AI orchestration frameworks (preferred).
  • Experience with vector databases such as Qdrant, Pinecone, Weaviate, or Databricks Vector Search (preferred).

Responsibilities

  • Design and build AI/LLM data pipelines supporting training, inference, evaluation, embeddings, and retrieval workloads.
  • Build production-grade RAG systems covering ingestion, chunking, embedding generation, indexing, retrieval, reranking, and context construction.
  • Develop AI applications using LLMs, structured outputs, function and tool calling, and agentic workflows.
  • Build and optimize semantic search and vector retrieval systems.
  • Develop frameworks for LLM evaluation, monitoring, tracing, quality measurement, latency analysis, and cost optimization.
  • Design scalable batch and streaming pipelines using Databricks, Apache Spark, Delta Lake, Snowflake, and Airflow.
  • Establish data quality, governance, lineage, security, and observability practices.
  • Partner with ML and application engineering teams to transition AI prototypes into reliable, production-ready systems.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now