Staff Data Engineer, Data Platform

New
D
DuettoHospitality Tech
Remote (US)Full-TimeStaff
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
7+ years
Required Skills
PythonSQLAirflowCI/CDPySpark

Requirements

  • 7+ years building production data systems in Python.
  • Deep expertise in PySpark and distributed data processing (Glue, EMR, or Databricks).
  • Strong experience with lakehouse architectures: Iceberg, Delta Lake, or Hudi on S3.
  • Production experience with Airflow or a comparable workflow orchestrator.
  • Solid AWS production experience across S3, Glue, Athena, Lambda, and SQS.
  • Track record of improving data quality, governance, and pipeline reliability at scale.
  • Working knowledge of Java for reading upstream systems preferred.
  • Experience with Trino or Presto for interactive SQL analytics preferred.
  • Experience with dbt for data transformation and modeling preferred.
  • Familiarity with Great Expectations or similar data quality frameworks preferred.
  • Genuine interest in AI-assisted development and LLM-based tooling.

Responsibilities

  • Own the design, performance, and reliability of Duetto's data lakehouse including bronze to gold architecture evolution.
  • Architect the shift from batch to near-real-time streaming using SQS-driven pipelines and Iceberg sinks.
  • Drive data quality and governance at scale by leading adoption of data contracts and formalizing schemas.
  • Strengthen observability and reliability using Datadog, Sentry, and Sumo Logic.
  • Build and maintain shared internal Python libraries and optimize CI/CD deployment workflows.
  • Contribute to AI-assisted pipeline generation and automated data quality using specialized LLM-based agent systems.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now