AWS Lakehouse Data Engineer

New
D
Delan Associates, IncData engineering
100% RemoteContractMiddle
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
SIX (6) years of relevant experience.
Required Skills
AWSPythonSQLCI/CDPySpark

Requirements

  • Hold a bachelor's degree in Engineering, Information Technology, Computer Science, Data Engineering, or a related field, or have four years of equivalent practical experience in lieu of a degree.
  • Have six years of relevant experience.
  • Have hands-on experience implementing AWS-native data lake or lakehouse architectures using Amazon S3 and services such as AWS Glue, Amazon Athena, Amazon EMR, AWS Lake Formation, and Amazon Redshift.
  • Have strong experience developing production ETL/ELT pipelines using Python and PySpark, including data modeling, transformation, testing, performance tuning, and error handling.
  • Have hands-on Apache Iceberg experience, including ACID transactions, snapshots, schema and partition evolution, time travel, table maintenance, and query optimization.
  • Have advanced SQL skills and experience supporting analytical queries, semantic layers, reporting tools, and data visualization workloads.
  • Have experience implementing metadata management and governance, including cataloging, lineage, ownership, classification, policy enforcement, and fine-grained access controls.
  • Have experience with AWS security fundamentals, including IAM, least privilege, KMS encryption, secrets management, network security, logging, and secure SDLC practices.
  • Have experience provisioning AWS resources using IaC and operating data platforms across multiple environments.
  • Have experience building or operating CI/CD pipelines for data workflows, including testing, packaging, deployment automation, environment promotion, and rollback.
  • Be able to troubleshoot distributed data-processing workloads and optimize performance, reliability, and cost.

Responsibilities

  • Design and implement batch and streaming ingestion from APIs, relational databases, file drops, event streams, and external partners.
  • Build, test, and optimize Python and PySpark ETL/ELT pipelines and curated datasets for analytics, reporting, visualization, and machine learning.
  • Build and operate an AWS-native lakehouse on Amazon S3 using Apache Iceberg and open columnar formats such as Apache Parquet.
  • Enable querying with AWS services such as Amazon Athena, Amazon EMR, AWS Glue, and Amazon Redshift.
  • Implement metadata management, data lineage, governance, and fine-grained access controls using AWS-native services.
  • Build data quality checks and publish measurable SLAs/SLOs.
  • Automate AWS provisioning and CI/CD for data pipelines and lakehouse components across environments.
  • Implement observability, operational runbooks, and incident-response procedures, and improve platform performance, reliability, security, and cost.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now