Engenheiro de Dados Pleno
J
JobgetherData engineering, AI
Based in BrazilFull-TimeMiddle
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- PythonSQLCI/CDData modelingDatabricksPySpark
Requirements
- Have advanced Python and SQL skills.
- Have solid knowledge of data modeling and medallion/lakehouse architecture.
- Have experience building pipelines for unstructured data, including PDFs, images, and text, and preparing datasets for RAG applications.
- Be familiar with vector databases and embeddings, such as Databricks Vector Search, pgvector, or similar technologies.
- Have experience with data quality, data testing, and versioning practices.
- Have hands-on Databricks experience, including Delta Lake, Workflows/Jobs, notebooks, PySpark, and Spark SQL.
- Have experience with Git and CI/CD practices.
- Have professional experience working in cloud environments such as Azure, AWS, or GCP.
- Know LGPD and appropriate handling of sensitive data.
- Unity Catalog experience in governed environments is desirable.
- Familiarity with MLflow, Databricks Model Serving, or Mosaic AI is a plus.
- Experience with Delta Live Tables, Lakeflow, or Auto Loader is desirable.
Responsibilities
- Design and build ETL/ELT pipelines using Databricks or open-source technologies and Delta Lake's bronze, silver, and gold medallion architecture.
- Ingest, standardize, and enrich unstructured documents such as PDFs and images with metadata.
- Curate and anonymize data in accordance with LGPD requirements for RAG, few-shot workflows, and testing.
- Build and operate vector indexing pipelines, including chunking, embedding generation, incremental updates, and version control.
- Prepare versioned datasets and golden sets for evaluation frameworks and accuracy benchmarking.
- Structure persistence for feedback cycles and support quality and SLA dashboards.
- Integrate data platforms with inference services, APIs, and other systems, considering performance, cost, monitoring, and alerting.
- Automate data workflows through Jobs and Workflows, CI/CD pipelines, and data quality testing.
- Improve data engineering practices, reliability, and operational efficiency.
View Full Description & ApplyYou'll be redirected to the employer's site