AWS Lakehouse Data Engineer
New
D
Delan Associates, IncData engineering
100% RemoteContractMiddle
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- SIX (6) years of relevant experience.
- Required Skills
- AWSPythonSQLCI/CDPySpark
Requirements
- Hold a bachelor's degree in Engineering, Information Technology, Computer Science, Data Engineering, or a related field, or have four years of equivalent practical experience in lieu of a degree.
- Have six years of relevant experience.
- Have hands-on experience implementing AWS-native data lake or lakehouse architectures using Amazon S3 and services such as AWS Glue, Amazon Athena, Amazon EMR, AWS Lake Formation, and Amazon Redshift.
- Have strong experience developing production ETL/ELT pipelines using Python and PySpark, including data modeling, transformation, testing, performance tuning, and error handling.
- Have hands-on Apache Iceberg experience, including ACID transactions, snapshots, schema and partition evolution, time travel, table maintenance, and query optimization.
- Have advanced SQL skills and experience supporting analytical queries, semantic layers, reporting tools, and data visualization workloads.
- Have experience implementing metadata management and governance, including cataloging, lineage, ownership, classification, policy enforcement, and fine-grained access controls.
- Have experience with AWS security fundamentals, including IAM, least privilege, KMS encryption, secrets management, network security, logging, and secure SDLC practices.
- Have experience provisioning AWS resources using IaC and operating data platforms across multiple environments.
- Have experience building or operating CI/CD pipelines for data workflows, including testing, packaging, deployment automation, environment promotion, and rollback.
- Be able to troubleshoot distributed data-processing workloads and optimize performance, reliability, and cost.
Responsibilities
- Design and implement batch and streaming ingestion from APIs, relational databases, file drops, event streams, and external partners.
- Build, test, and optimize Python and PySpark ETL/ELT pipelines and curated datasets for analytics, reporting, visualization, and machine learning.
- Build and operate an AWS-native lakehouse on Amazon S3 using Apache Iceberg and open columnar formats such as Apache Parquet.
- Enable querying with AWS services such as Amazon Athena, Amazon EMR, AWS Glue, and Amazon Redshift.
- Implement metadata management, data lineage, governance, and fine-grained access controls using AWS-native services.
- Build data quality checks and publish measurable SLAs/SLOs.
- Automate AWS provisioning and CI/CD for data pipelines and lakehouse components across environments.
- Implement observability, operational runbooks, and incident-response procedures, and improve platform performance, reliability, security, and cost.
View Full Description & ApplyYou'll be redirected to the employer's site