Senior Data QA Automation Engineer
New
V
v4c.aiData Engineering
United StatesFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 8+ years of experience in data engineering, data QA, or software development engineering in test (SDET), with at least 2+ years of dedicated experience architecting test automation in Databricks
- Required Skills
- PythonSQLSparkCI/CDDatabricksPySpark
Requirements
- Bachelor’s or Master’s degree in Computer Science, Data Engineering, or a related quantitative field.
- 8+ years of experience in data engineering, data QA, or software development engineering in test (SDET).
- 2+ years of dedicated experience architecting test automation in Databricks.
- Mastery of Python and PySpark (DataFrames and SQL APIs) for processing and profiling large datasets.
- Deep expertise in writing advanced SQL queries, optimization techniques, and understanding Spark query execution plans.
- Hands-on mastery of big-data validation libraries such as Great Expectations, pytest, or Delta Live Tables expectations.
- Strong operational knowledge of Databricks deployment on a major cloud provider (AWS, Azure, or GCP).
- Experience validating real-time event-streaming architectures (Kafka, Event Hubs, Kinesis) is preferred.
- Databricks Certified Data Engineer Professional or Machine Learning Professional certifications are preferred.
Responsibilities
- Architect, build, and scale automated test frameworks from scratch natively within Databricks using PySpark, Python, and SQL.
- Design robust automated assertions for Delta Lake tables, including checking data drift, schema evolution, and historical data validation via time-travel functions.
- Code complex automated scenarios to validate large-scale batch and real-time streaming data pipelines, ensuring source-to-target integrity.
- Programmatically verify data lineage, audit logs, and access controls implemented via Databricks Unity Catalog.
- Lead the integration of automated data quality tests into enterprise CI/CD pipelines leveraging Databricks Workflows, APIs, or Airflow.
- Act as the subject matter expert for data quality; mentor junior team members, establish QA standards, and advocate for data quality principles.
- Design and execute automated performance and scalability tests on Spark jobs, large clusters, and complex query optimizations.
View Full Description & ApplyYou'll be redirected to the employer's site