Engenheiro de Dados Pleno

J
JobgetherData engineering, AI
Based in BrazilFull-TimeMiddle
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Required Skills
PythonSQLCI/CDData modelingDatabricksPySpark

Requirements

  • Have advanced Python and SQL skills.
  • Have solid knowledge of data modeling and medallion/lakehouse architecture.
  • Have experience building pipelines for unstructured data, including PDFs, images, and text, and preparing datasets for RAG applications.
  • Be familiar with vector databases and embeddings, such as Databricks Vector Search, pgvector, or similar technologies.
  • Have experience with data quality, data testing, and versioning practices.
  • Have hands-on Databricks experience, including Delta Lake, Workflows/Jobs, notebooks, PySpark, and Spark SQL.
  • Have experience with Git and CI/CD practices.
  • Have professional experience working in cloud environments such as Azure, AWS, or GCP.
  • Know LGPD and appropriate handling of sensitive data.
  • Unity Catalog experience in governed environments is desirable.
  • Familiarity with MLflow, Databricks Model Serving, or Mosaic AI is a plus.
  • Experience with Delta Live Tables, Lakeflow, or Auto Loader is desirable.

Responsibilities

  • Design and build ETL/ELT pipelines using Databricks or open-source technologies and Delta Lake's bronze, silver, and gold medallion architecture.
  • Ingest, standardize, and enrich unstructured documents such as PDFs and images with metadata.
  • Curate and anonymize data in accordance with LGPD requirements for RAG, few-shot workflows, and testing.
  • Build and operate vector indexing pipelines, including chunking, embedding generation, incremental updates, and version control.
  • Prepare versioned datasets and golden sets for evaluation frameworks and accuracy benchmarking.
  • Structure persistence for feedback cycles and support quality and SLA dashboards.
  • Integrate data platforms with inference services, APIs, and other systems, considering performance, cost, monitoring, and alerting.
  • Automate data workflows through Jobs and Workflows, CI/CD pipelines, and data quality testing.
  • Improve data engineering practices, reliability, and operational efficiency.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now