Senior Data Engineer
New
Z
ZifoClinical trial data
Workable workplace: remote; Workable locations: United StatesFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- AWSPythonFastAPI
Requirements
- Strong Python development experience, including PDF or document parsing libraries such as PyMuPDF, pdfplumber, or unstructured.io.
- Advanced PostgreSQL-compatible SQL experience, including schema design, migrations, query optimization, and indexing; Amazon Aurora experience is specified.
- Hands-on experience with Neptune, Neo4j, or similar graph databases, and proficiency in SPARQL or Cypher.
- Experience with ontology and knowledge graph modeling for biomedical entities.
- Experience with AWS services including Aurora PostgreSQL, S3, Lambda, Step Functions, SQS/SNS, and IAM.
- Experience with PDF text extraction, document section classification, and named entity recognition for clinical or biomedical text.
- Familiarity with embedding models and vector stores such as OpenSearch, pgvector, or Pinecone.
- Experience building data-serving APIs using FastAPI, including asynchronous programming patterns and backend integration.
- Experience preparing data for LangChain or LangGraph applications and designing RAG pipelines.
- Experience with workflow orchestration tools such as Airflow, Prefect, AWS Step Functions, or Temporal.
- Experience with Terraform or AWS CDK, Docker, and Git, including automated pipeline testing and AWS deployment.
- Understanding of clinical trial structure and familiarity with clinical data standards or terminologies such as MeSH, MedDRA, SNOMED, ATC codes, or CDISC.
Responsibilities
- Build ingestion pipelines for clinical trial protocols, ICF documents, SmPCs, CSRs, and published articles.
- Parse PDFs, extract text, and normalize unstructured clinical content into structured data.
- Design data models in Amazon Aurora and GraphDB for trial design entities and their relationships.
- Develop embedding and vectorization pipelines for RAG-based retrieval in LangGraph workflows.
- Build and maintain ETL/ELT workflows across relational and graph data stores.
- Implement clinical data quality validation, including section classification, entity completeness, and cross-reference integrity.
- Build data-serving APIs with Python and FastAPI for the Angular frontend and LangGraph agent layer.
- Set up data lineage tracking and audit trails for regulatory traceability.
View Full Description & ApplyYou'll be redirected to the employer's site