Director Analytics Infrastructure, Pipeline Operations
New
N
NovartisLife Sciences
Remote Position (USA), Remote, USFull-TimeDirector
SalaryUSD 194600 - 361400 / year
Apply NowOpens the employer's application page
Job Details
- Languages
- English
- Experience
- 10+ years of experience in data engineering, ML/AI engineering, or analytics infrastructure; 5+ years leading teams building enterprise-scale data platforms and feature stores.
- Required Skills
- PythonSQLSnowflakeAirflowSparkdbtDatabricksMLOpsPySpark
Requirements
- Advanced degree in Computer Science, Data Engineering, or related field.
- 10+ years of experience in data engineering, ML/AI engineering, or analytics infrastructure.
- 5+ years leading teams building enterprise-scale data platforms and feature stores.
- Expert knowledge of feature store technologies (Feast, Tecton, SageMaker Feature Store, Databricks Feature Store).
- Deep expertise in modern data platforms optimized for ML workloads (Databricks, Auto ML, Snowflake, BigQuery).
- Strong proficiency in Python, SQL, Spark/PySpark for large-scale data processing.
- Experience with data orchestration tools (Airflow, Prefect, dbt) and CI/CD for data pipelines.
- Understanding of data governance, privacy (HIPAA, GDPR), and compliance in life sciences.
- 20% travel required.
Responsibilities
- Design and implement intelligent, self-healing data pipelines that leverage AI/ML for automated data quality monitoring, anomaly detection, and remediation.
- Build and maintain centralized feature stores that enable feature reusability across multiple models and use cases.
- Create curated data repositories optimized for data science/AI workflows, including training datasets, evaluation datasets, and production serving layers.
- Develop automated feature engineering pipelines that transform raw data into analytics-ready features with lineage tracking.
- Partner with Enterprise IT to optimize analytics platform architecture for high-performance data science workloads.
- Build automated pipelines that integrate diverse data sources including sales, CRM, patient claims, real-world evidence, and unstructured data.
- Create self-service data access layers that empower data scientists and analysts to query and extract data independently.
- Establish SLAs for data availability, freshness, and quality; implement monitoring and observability solutions.
View Full Description & ApplyYou'll be redirected to the employer's site