AWS Lakehouse Data Engineer
New
I
Inizio Partners CorpData engineering
US - Remote (Any location)ContractSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- SIX (6) years of relevant experience.
- Required Skills
- PythonSQLPySpark
Requirements
- Have a bachelor's degree in Engineering, Information Technology, Computer Science, Data Engineering, or a related field, or four years of equivalent practical experience in lieu of a degree.
- Have six years of relevant experience.
- Have hands-on experience implementing AWS-native data lake or lakehouse architectures using Amazon S3 and services such as AWS Glue, Amazon Athena, Amazon EMR, AWS Lake Formation, and Amazon Redshift.
- Have production ETL/ELT pipeline experience with Python and PySpark, including data modeling, transformation, testing, performance tuning, and error handling.
- Have hands-on Apache Iceberg experience, including ACID transactions, snapshots, schema and partition evolution, time travel, table maintenance, and query optimization.
- Have advanced SQL skills and experience supporting analytical queries, semantic layers, reporting tools, and data visualization workloads.
- Have experience with metadata management and governance, including cataloging, lineage, ownership, classification, policy enforcement, and fine-grained access controls.
- Have experience with AWS security fundamentals, including IAM, least privilege, KMS encryption, secrets management, network security, logging, and secure SDLC practices.
- Have experience provisioning AWS resources using IaC and operating data platforms across multiple environments.
- Have experience building or operating CI/CD pipelines for data workflows, including testing, packaging, deployment automation, environment promotion, and rollback.
- Be able to troubleshoot distributed data-processing workloads and optimize performance, reliability, and cost.
- Be able to obtain Public Trust.
Responsibilities
- Design and implement batch and streaming ingestion from APIs, relational databases, file drops, event streams, and external partners.
- Develop, test, and optimize Python and PySpark ETL/ELT pipelines, including incremental processing, CDC, data contracts, and schema validation.
- Build and operate an AWS-native lakehouse on Amazon S3 using Apache Iceberg and open columnar formats such as Apache Parquet.
- Enable lakehouse querying with AWS services such as Amazon Athena, Amazon EMR, AWS Glue, and Amazon Redshift.
- Optimize data platform performance and cost through partitioning, compaction, file sizing, statistics, caching, and lifecycle policies.
- Implement metadata management, governance, fine-grained access controls, and end-to-end data lineage.
- Build data quality checks and publish measurable SLAs and SLOs.
- Automate AWS provisioning and CI/CD, including tests, security checks, deployment, promotion, and rollback.
- Implement observability using metrics, logs, traces, alerts, dashboards, runbooks, and incident-response procedures.
- Collaborate with data, application, analytics, AI/ML, security, networking, and cloud platform teams; maintain engineering documentation.
View Full Description & ApplyYou'll be redirected to the employer's site