Data Engineer
New
V
Vaniam GroupHealthcare data
Location: United StatesFull-TimeMiddle
Salary110,000 - 125,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years of professional experience in data engineering, ETL, or related roles
- Required Skills
- DockerPythonSQLAirflowSparkdbtPySpark
Requirements
- Have 5+ years of professional experience in data engineering, ETL, or related roles.
- Demonstrate strong proficiency in Python and SQL for data engineering.
- Have hands-on experience building and maintaining pipelines in a lakehouse or modern data platform.
- Understand Medallion architectures and layered data design.
- Be familiar with Spark or PySpark.
- Have familiarity with workflow orchestration tools such as Airflow or dbt, or similar tools.
- Be familiar with testing and observability frameworks.
- Have experience with Docker and Git-based version control.
- Experience with Databricks and the Microsoft Azure ecosystem is a plus.
- Experience with Delta Lake formats, metadata management, or data catalogs is a plus.
- Familiarity with healthcare, scientific, or engagement data domains is a plus.
- Experience exposing analytics through APIs or lightweight microservices is a plus.
Responsibilities
- Design, build, and operate reliable ETL and ELT pipelines in Python and SQL.
- Manage ingestion, standardization, quality, and curated serving across Bronze, Silver, and Gold layers.
- Maintain ingestion from transactional MySQL systems into Vaniam Core.
- Implement observability, data quality checks, and lineage tracking for downstream datasets.
- Develop schemas, tables, and views for analytics, APIs, and product use cases.
- Apply security, privacy, compliance, and access-control practices across sensitive healthcare data.
- Integrate third-party, client-provided, and product-generated data sources into Vaniam Core.
- Collaborate with innovation, product, Data Science, and AI teams to productionize analytics and predictive tools.
- Monitor job execution, storage, and cluster performance; troubleshoot issues and address bottlenecks.
- Conduct code reviews, enforce standards, and contribute to CI/CD practices for data pipelines.
View Full Description & ApplyYou'll be redirected to the employer's site