- Design, build, and maintain high-throughput ETL/ELT pipelines for data ingestion, processing, and storage solutions.
- Develop complex, performance-tuned data transformations using Python, high-performance SQL, and tools like dbt.
- Implement data ops best practices, including CI/CD for data pipelines, version-controlled schemas, and automated testing.
- Work with distributed processing systems like Spark, Flink, or Kafka to support scalable batch and real-time operational analytics.
- Partner with the ML Platform team to prepare and provide clean, feature-rich datasets for model training and inference.
- Ensure data stewardship by contributing to documentation, discoverability, and implementing robust data privacy and access controls.
- Collaborate with Accounting, FP&A, and Finance Ops to translate close workflows and audit evidence needs into concrete data requirements.