- Own data pipeline design and implementation; consolidate 250+ transactional tables into ~10 analytic tables.
- Transition from weekly full-refresh to daily/near-real-time incremental loads using CDC tools like AWS DMS.
- Build and own a centralized analytics layer to serve BI, product, and AI needs.
- Establish unified data definitions, metrics, KPIs, and a Master Data Management (MDM) framework with governance standards.
- Partner with the Applied AI Engineer to design AI-ready data models including structures for RAG and LLM applications.
- Drive migration to the centralized layer while deprecating direct raw table access.
- Partner with internal departments to drive self-service analytics and build data marts for benchmarking.
- Conduct exploratory analysis for data validation and implement observability and monitoring tools.
- Expand the data lake to ingest third-party sources including HubSpot, Jira, and Zendesk.
AWSPostgreSQLPython+3 more