Director Analytics Infrastructure, Pipeline Operations

New
N
NovartisLife Sciences
Remote Position (USA), Remote, USFull-TimeDirector
SalaryUSD 194600 - 361400 / year
Apply NowOpens the employer's application page

Job Details

Languages
English
Experience
10+ years of experience in data engineering, ML/AI engineering, or analytics infrastructure; 5+ years leading teams building enterprise-scale data platforms and feature stores.
Required Skills
PythonSQLSnowflakeAirflowSparkdbtDatabricksMLOpsPySpark

Requirements

  • Advanced degree in Computer Science, Data Engineering, or related field.
  • 10+ years of experience in data engineering, ML/AI engineering, or analytics infrastructure.
  • 5+ years leading teams building enterprise-scale data platforms and feature stores.
  • Expert knowledge of feature store technologies (Feast, Tecton, SageMaker Feature Store, Databricks Feature Store).
  • Deep expertise in modern data platforms optimized for ML workloads (Databricks, Auto ML, Snowflake, BigQuery).
  • Strong proficiency in Python, SQL, Spark/PySpark for large-scale data processing.
  • Experience with data orchestration tools (Airflow, Prefect, dbt) and CI/CD for data pipelines.
  • Understanding of data governance, privacy (HIPAA, GDPR), and compliance in life sciences.
  • 20% travel required.

Responsibilities

  • Design and implement intelligent, self-healing data pipelines that leverage AI/ML for automated data quality monitoring, anomaly detection, and remediation.
  • Build and maintain centralized feature stores that enable feature reusability across multiple models and use cases.
  • Create curated data repositories optimized for data science/AI workflows, including training datasets, evaluation datasets, and production serving layers.
  • Develop automated feature engineering pipelines that transform raw data into analytics-ready features with lineage tracking.
  • Partner with Enterprise IT to optimize analytics platform architecture for high-performance data science workloads.
  • Build automated pipelines that integrate diverse data sources including sales, CRM, patient claims, real-world evidence, and unstructured data.
  • Create self-service data access layers that empower data scientists and analysts to query and extract data independently.
  • Establish SLAs for data availability, freshness, and quality; implement monitoring and observability solutions.
View Full Description & ApplyYou'll be redirected to the employer's site
USD 194600 - 361400 / year
Apply Now