Data Engineering Specialist – AI

New
J
JobgetherArtificial Intelligence
United StatesFull-TimeSenior
Salary100,000 - 150,000 USD per year
Apply NowOpens the employer's application page

Job Details

Experience
6+ years
Required Skills
PythonMachine LearningSparkCI/CDData modelingDistributed Systems

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field.
  • 6+ years of professional experience in data engineering, including significant experience supporting AI or machine learning environments.
  • Strong programming skills in Python and experience with at least one JVM-based or systems programming language.
  • Hands-on experience with large-scale data processing frameworks such as Spark, Ray, or Beam.
  • Experience operating large-scale storage systems and high-volume data pipelines.
  • Strong understanding of distributed systems, data modeling, and modern storage formats.
  • Experience implementing dataset versioning, lineage, and reproducibility solutions for machine learning workflows.
  • Familiarity with high-performance data loading architectures for accelerator-based model training.
  • Strong software engineering practices, including testing, CI/CD workflows, and code review processes.
  • Excellent problem-solving, communication, and cross-functional collaboration skills.

Responsibilities

  • Design and operate large-scale data pipelines supporting AI training, evaluation, and continuous improvement processes.
  • Build ingestion systems capable of handling diverse data modalities, including text, images, audio, video, and structured datasets.
  • Develop data cleaning, deduplication, filtering, and quality assurance processes at significant scale.
  • Implement dataset versioning, lineage tracking, and provenance systems to ensure reproducible AI workflows.
  • Build high-throughput data loading solutions that optimize accelerator and GPU utilization during training.
  • Develop labeling workflows, active learning systems, and human-in-the-loop processes to improve dataset quality.
  • Design storage architectures that balance cost, scalability, throughput, and performance requirements.
  • Create evaluation dataset pipelines with strong integrity controls and contamination prevention measures.
  • Implement privacy, security, redaction, and consent mechanisms throughout data workflows.
View Full Description & ApplyYou'll be redirected to the employer's site
100,000 - 150,000 USD per year
Apply Now