Senior Data Engineer – Clinical Platforms (Databricks)

New
M
MuttdataClinical software
Remote - LatamFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Languages
Advanced English
Experience
4+ years of hands-on experience building production data pipelines and lakehouse architectures using Databricks, Delta Lake, and Apache Spark (PySpark or Scala).
Required Skills
PythonSQLSparkRESTful APIsScalaDatabricksPySpark

Requirements

  • Have 4+ years of hands-on experience building production data pipelines and lakehouse architectures using Databricks, Delta Lake, and Apache Spark with PySpark or Scala.
  • Have experience building, extending, or maintaining custom clinical trial software applications, such as custom EDC, CTMS, Clinical Data Repositories, or eCOA/ePRO platforms.
  • Understand clinical data standards and regulatory environments, including CDISC (SDTM, ADaM, CDASH), 21 CFR Part 11, GxP validation, and ICH-GCP guidelines.
  • Have experience with relational schema design, dimensional modeling, and unstructured data handling within Delta Lake environments.
  • Be proficient in Python, SQL, RESTful API integrations, CI/CD pipelines, Git, and automated testing frameworks.
  • Have experience working in cloud environments; AWS is preferred, or Azure or GCP.
  • Have advanced English to discuss technical requirements and solutions with clients in the United States.
  • A bachelor's or master's degree in Computer Science, Data Engineering, Bioinformatics, or a related quantitative field is a plus.
  • Experience with Databricks Workflows, Delta Live Tables (DLT), and Unity Catalog governance is a plus.
  • A background working in a validated system environment (Computer System Validation / CSV) is a plus.

Responsibilities

  • Design, build, and optimize enterprise data pipelines, lakehouse storage layers, and data models using Databricks, PySpark, Spark SQL, and Delta Lake.
  • Collaborate with frontend developers, software architects, and clinical research teams to build API-driven endpoints, data ingestion engines, and query layers.
  • Build data structures for EDC outputs, audit trails, device telemetry, and patient-reported outcomes.
  • Partner with Clinical QA and Validation teams to support compliance with GxP, 21 CFR Part 11, HIPAA, and GDPR.
  • Implement real-time and batch ingestion jobs connecting legacy clinical systems, central labs, EHRs, and wearable devices.
  • Monitor, troubleshoot, and optimize Spark jobs, Delta Lake tables, and query execution times.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now