Data Engineering Specialist – AI
New
J
JobgetherArtificial Intelligence
United StatesFull-TimeSenior
Salary100,000 - 150,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 6+ years
- Required Skills
- PythonMachine LearningSparkCI/CDData modelingDistributed Systems
Requirements
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field.
- 6+ years of professional experience in data engineering, including significant experience supporting AI or machine learning environments.
- Strong programming skills in Python and experience with at least one JVM-based or systems programming language.
- Hands-on experience with large-scale data processing frameworks such as Spark, Ray, or Beam.
- Experience operating large-scale storage systems and high-volume data pipelines.
- Strong understanding of distributed systems, data modeling, and modern storage formats.
- Experience implementing dataset versioning, lineage, and reproducibility solutions for machine learning workflows.
- Familiarity with high-performance data loading architectures for accelerator-based model training.
- Strong software engineering practices, including testing, CI/CD workflows, and code review processes.
- Excellent problem-solving, communication, and cross-functional collaboration skills.
Responsibilities
- Design and operate large-scale data pipelines supporting AI training, evaluation, and continuous improvement processes.
- Build ingestion systems capable of handling diverse data modalities, including text, images, audio, video, and structured datasets.
- Develop data cleaning, deduplication, filtering, and quality assurance processes at significant scale.
- Implement dataset versioning, lineage tracking, and provenance systems to ensure reproducible AI workflows.
- Build high-throughput data loading solutions that optimize accelerator and GPU utilization during training.
- Develop labeling workflows, active learning systems, and human-in-the-loop processes to improve dataset quality.
- Design storage architectures that balance cost, scalability, throughput, and performance requirements.
- Create evaluation dataset pipelines with strong integrity controls and contamination prevention measures.
- Implement privacy, security, redaction, and consent mechanisms throughout data workflows.
View Full Description & ApplyYou'll be redirected to the employer's site