Data Engineer (Generative AI & Healthcare Data)
New
A
AltarumHealthcare data and AI
All work must be performed within the continental U.S. for the duration of employment, unless required by contract., Ability to work core hours aligned with Eastern Time, unless otherwise approved by your manager.Full-TimeMiddle
Salary144,771 - 174,723.15 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- Typically 3–5 years of relevant experience
- Required Skills
- PythonSQLSparkCI/CDDatabricksGenerative AILangChain
Requirements
- Bachelor's degree in computer science, data science, information systems, engineering, mathematics, statistics, public health informatics, or a related field, or equivalent practical experience.
- Typically 3–5 years of relevant experience in data engineering, AI engineering, software engineering, analytics engineering, cloud engineering, or applied technical research, including independent delivery of production systems.
- Strong Python and SQL skills, including building maintainable code, data models, automated tests, and reliable ingestion and transformation pipelines.
- Hands-on experience with a cloud platform and a modern data platform such as Databricks or an equivalent, including distributed processing with Spark or comparable technology.
- Experience using Git, CI/CD, environment-specific configuration, and automated deployment for production workloads.
- Hands-on experience building and deploying generative AI applications with LangChain, LlamaIndex, or comparable frameworks.
- Experience with retrieval-augmented generation or document processing, prompt design, API integration, evaluation, and monitoring.
- Ability to turn requirements into technical designs, manage defined work, troubleshoot independently, and guide collaborators.
- Databricks Certified Generative AI Engineer Associate certification is strongly preferred; comparable generative AI credentials with relevant practical experience are also valued.
- Preferred experience includes Databricks Mosaic AI, Vector Search, Model Serving, MLflow evaluation and tracing, Unity Catalog, healthcare and Medicaid data, FHIR or HL7, Delta Lake, PySpark, dbt, orchestration, or infrastructure as code.
Responsibilities
- Design and operate cloud data platform components, including lakehouse storage, compute, SQL and ELT engines, and reusable transformation workflows.
- Develop production Python and SQL code, reusable packages, APIs, and configuration-driven frameworks with automated tests and documentation.
- Own CI/CD, deployment, release verification, and rollback for data pipelines, AI applications, and supporting services.
- Build ingestion, transformation, and normalization pipelines for structured, semi-structured, and unstructured healthcare and public health data.
- Integrate healthcare data such as Medicaid and Medicare claims, FHIR R4, US Core, HL7 v2, and social determinants of health.
- Define data quality, validation, lineage, metadata, monitoring, and reliability practices for owned pipelines and services.
- Build retrieval-augmented generation and document-intelligence applications, including ingestion, chunking, embeddings, vector search, and retrieval.
- Evaluate, deploy, monitor, and support AI applications and model endpoints, managing output quality, factual accuracy, latency, and cost.
- Apply privacy, security, and responsible AI controls, and document model behavior, fairness, explainability, and governance.
- Lead defined project workstreams, review implementation choices, mentor colleagues, and create technical guidance and learning materials.
View Full Description & ApplyYou'll be redirected to the employer's site