Senior Data Engineer

New
Z
ZifoClinical trial data
Workable workplace: remote; Workable locations: United StatesFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Required Skills
AWSPythonFastAPI

Requirements

  • Strong Python development experience, including PDF or document parsing libraries such as PyMuPDF, pdfplumber, or unstructured.io.
  • Advanced PostgreSQL-compatible SQL experience, including schema design, migrations, query optimization, and indexing; Amazon Aurora experience is specified.
  • Hands-on experience with Neptune, Neo4j, or similar graph databases, and proficiency in SPARQL or Cypher.
  • Experience with ontology and knowledge graph modeling for biomedical entities.
  • Experience with AWS services including Aurora PostgreSQL, S3, Lambda, Step Functions, SQS/SNS, and IAM.
  • Experience with PDF text extraction, document section classification, and named entity recognition for clinical or biomedical text.
  • Familiarity with embedding models and vector stores such as OpenSearch, pgvector, or Pinecone.
  • Experience building data-serving APIs using FastAPI, including asynchronous programming patterns and backend integration.
  • Experience preparing data for LangChain or LangGraph applications and designing RAG pipelines.
  • Experience with workflow orchestration tools such as Airflow, Prefect, AWS Step Functions, or Temporal.
  • Experience with Terraform or AWS CDK, Docker, and Git, including automated pipeline testing and AWS deployment.
  • Understanding of clinical trial structure and familiarity with clinical data standards or terminologies such as MeSH, MedDRA, SNOMED, ATC codes, or CDISC.

Responsibilities

  • Build ingestion pipelines for clinical trial protocols, ICF documents, SmPCs, CSRs, and published articles.
  • Parse PDFs, extract text, and normalize unstructured clinical content into structured data.
  • Design data models in Amazon Aurora and GraphDB for trial design entities and their relationships.
  • Develop embedding and vectorization pipelines for RAG-based retrieval in LangGraph workflows.
  • Build and maintain ETL/ELT workflows across relational and graph data stores.
  • Implement clinical data quality validation, including section classification, entity completeness, and cross-reference integrity.
  • Build data-serving APIs with Python and FastAPI for the Angular frontend and LangGraph agent layer.
  • Set up data lineage tracking and audit trails for regulatory traceability.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now