Software Development Engineer II – Data Engineer
New
G
Gather AISupply chain robotics
Remote (India); fully remote role on our India-based team., We work across multiple time zones.Full-TimeMiddle
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Languages
- Clear written and spoken English.
- Experience
- 2–5 years building and running production data pipelines
- Required Skills
- PostgreSQLPythonSQLAirflowAzureCI/CDdbt
Requirements
- Have 2–5 years of experience building and running production data pipelines.
- Hold a degree in Computer Science or have equivalent practical experience.
- Use strong SQL, including joins, window functions, and CTEs, and understand dimensional modelling, including facts, dimensions, and grain.
- Have hands-on experience building tested models in dbt, Snowflake Dynamic Tables, Databricks Lakeflow Declarative Pipelines, or an equivalent tool.
- Write production Python pipeline code with tests and code review, beyond notebooks or one-off scripts.
- Have experience running pipelines with Airflow, Dagster, Databricks Lakeflow Jobs, Snowflake Tasks, or equivalent, including failures, reruns, and backfills.
- Have loaded data from an operational database into a warehouse or lakehouse such as Snowflake or Databricks.
- Have production experience on Azure or another major cloud, including object storage, Git, CI/CD, and Docker/Kubernetes.
- Write tests alongside code and care about data correctness.
- Communicate clearly in written and spoken English; write PRs and documentation and raise blockers early.
Responsibilities
- Build incremental extraction pipelines that move data from production PostgreSQL into the analytical warehouse while protecting the production database.
- Write and maintain shared dbt models, including dimensions, reusable metric building blocks, and serving tables.
- Add metrics to the semantic layer so measures are defined consistently across dashboards.
- Add tests, freshness checks, and alerts to pipelines, and help run safe backfills.
- Apply access-control and tenant-separation patterns to each model and pipeline.
- Maintain lineage from metrics to source records and drone images.
- Partner with the integration team to validate incoming WMS data against agreed contracts.
- Document models and register them in the data catalog.
- Deliver through CI/CD and join on-call for owned pipelines.
View Full Description & ApplyYou'll be redirected to the employer's site