Staff Data Engineer, Data Platform
New
D
DuettoHospitality Tech
Remote (US)Full-TimeStaff
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 7+ years
- Required Skills
- PythonSQLAirflowCI/CDPySpark
Requirements
- 7+ years building production data systems in Python.
- Deep expertise in PySpark and distributed data processing (Glue, EMR, or Databricks).
- Strong experience with lakehouse architectures: Iceberg, Delta Lake, or Hudi on S3.
- Production experience with Airflow or a comparable workflow orchestrator.
- Solid AWS production experience across S3, Glue, Athena, Lambda, and SQS.
- Track record of improving data quality, governance, and pipeline reliability at scale.
- Working knowledge of Java for reading upstream systems preferred.
- Experience with Trino or Presto for interactive SQL analytics preferred.
- Experience with dbt for data transformation and modeling preferred.
- Familiarity with Great Expectations or similar data quality frameworks preferred.
- Genuine interest in AI-assisted development and LLM-based tooling.
Responsibilities
- Own the design, performance, and reliability of Duetto's data lakehouse including bronze to gold architecture evolution.
- Architect the shift from batch to near-real-time streaming using SQS-driven pipelines and Iceberg sinks.
- Drive data quality and governance at scale by leading adoption of data contracts and formalizing schemas.
- Strengthen observability and reliability using Datadog, Sentry, and Sumo Logic.
- Build and maintain shared internal Python libraries and optimize CI/CD deployment workflows.
- Contribute to AI-assisted pipeline generation and automated data quality using specialized LLM-based agent systems.
View Full Description & ApplyYou'll be redirected to the employer's site