AWS Lakehouse Data Engineer

New
I
Inizio Partners CorpData engineering
US - Remote (Any location)ContractSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
SIX (6) years of relevant experience.
Required Skills
PythonSQLPySpark

Requirements

  • Have a bachelor's degree in Engineering, Information Technology, Computer Science, Data Engineering, or a related field, or four years of equivalent practical experience in lieu of a degree.
  • Have six years of relevant experience.
  • Have hands-on experience implementing AWS-native data lake or lakehouse architectures using Amazon S3 and services such as AWS Glue, Amazon Athena, Amazon EMR, AWS Lake Formation, and Amazon Redshift.
  • Have production ETL/ELT pipeline experience with Python and PySpark, including data modeling, transformation, testing, performance tuning, and error handling.
  • Have hands-on Apache Iceberg experience, including ACID transactions, snapshots, schema and partition evolution, time travel, table maintenance, and query optimization.
  • Have advanced SQL skills and experience supporting analytical queries, semantic layers, reporting tools, and data visualization workloads.
  • Have experience with metadata management and governance, including cataloging, lineage, ownership, classification, policy enforcement, and fine-grained access controls.
  • Have experience with AWS security fundamentals, including IAM, least privilege, KMS encryption, secrets management, network security, logging, and secure SDLC practices.
  • Have experience provisioning AWS resources using IaC and operating data platforms across multiple environments.
  • Have experience building or operating CI/CD pipelines for data workflows, including testing, packaging, deployment automation, environment promotion, and rollback.
  • Be able to troubleshoot distributed data-processing workloads and optimize performance, reliability, and cost.
  • Be able to obtain Public Trust.

Responsibilities

  • Design and implement batch and streaming ingestion from APIs, relational databases, file drops, event streams, and external partners.
  • Develop, test, and optimize Python and PySpark ETL/ELT pipelines, including incremental processing, CDC, data contracts, and schema validation.
  • Build and operate an AWS-native lakehouse on Amazon S3 using Apache Iceberg and open columnar formats such as Apache Parquet.
  • Enable lakehouse querying with AWS services such as Amazon Athena, Amazon EMR, AWS Glue, and Amazon Redshift.
  • Optimize data platform performance and cost through partitioning, compaction, file sizing, statistics, caching, and lifecycle policies.
  • Implement metadata management, governance, fine-grained access controls, and end-to-end data lineage.
  • Build data quality checks and publish measurable SLAs and SLOs.
  • Automate AWS provisioning and CI/CD, including tests, security checks, deployment, promotion, and rollback.
  • Implement observability using metrics, logs, traces, alerts, dashboards, runbooks, and incident-response procedures.
  • Collaborate with data, application, analytics, AI/ML, security, networking, and cloud platform teams; maintain engineering documentation.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now