- Design, build, and optimize ETL pipelines that power analytics, data science, and ML workflows using tools such as Databricks, PySpark, and Airflow.
- Develop and maintain labeling and retraining pipelines for machine learning models, ensuring quality, reproducibility, and observability.
- Implement and support MLOps practices, including model versioning, CI/CD for ML, and model monitoring in production environments.
- Collaborate with data scientists to productionize and scale model training, inference, and evaluation pipelines.
- Contribute to the design and evolution of the data lakehouse, including schema design, partitioning strategies, and performance optimization.
- Document and communicate data architecture, lineage, and dependencies to ensure transparency and maintainability across teams.
- Champion data quality and governance, ensuring that datasets are accurate, well-structured, and compliant with organizational standards.
- Leverage infrastructure-as-code and containerization to build reproducible, maintainable environments.
- Participate in code reviews and continuous improvement of engineering best practices within the team.