Senior Data Engineer (AI/ML)
New
J
JobgetherSecurity & IT
100% remote position across India., Flexible collaboration across international teams and time zones.Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years of experience in data engineering, software engineering, distributed systems, or a related field.
- Required Skills
- PythonSQLJavaSnowflakeAirflowSparkScalaDatabricksGenerative AI
Requirements
- 5+ years of experience in data engineering, software engineering, distributed systems, or a related field.
- Strong programming skills in Python and/or Scala/Java, combined with advanced SQL capabilities.
- Hands-on experience with Databricks, Snowflake, Apache Spark, Delta Lake, and Airflow.
- Strong experience working with cloud-based data platforms and scalable data architectures.
- Practical experience building applications using LLMs or Generative AI.
- Strong understanding of RAG architectures, embeddings, vector databases, semantic search, and retrieval systems.
- Experience with real-time and streaming architectures using technologies such as Kafka or Spark Structured Streaming.
- Experience with LLM/AI evaluation frameworks, automated evaluations, experimentation, and quality metrics.
- Familiarity with AI observability and tracing, including production monitoring.
- Experience with LangGraph, LangChain, LlamaIndex, or similar AI orchestration frameworks (preferred).
- Experience with vector databases such as Qdrant, Pinecone, Weaviate, or Databricks Vector Search (preferred).
Responsibilities
- Design and build AI/LLM data pipelines supporting training, inference, evaluation, embeddings, and retrieval workloads.
- Build production-grade RAG systems covering ingestion, chunking, embedding generation, indexing, retrieval, reranking, and context construction.
- Develop AI applications using LLMs, structured outputs, function and tool calling, and agentic workflows.
- Build and optimize semantic search and vector retrieval systems.
- Develop frameworks for LLM evaluation, monitoring, tracing, quality measurement, latency analysis, and cost optimization.
- Design scalable batch and streaming pipelines using Databricks, Apache Spark, Delta Lake, Snowflake, and Airflow.
- Establish data quality, governance, lineage, security, and observability practices.
- Partner with ML and application engineering teams to transition AI prototypes into reliable, production-ready systems.
View Full Description & ApplyYou'll be redirected to the employer's site