Data Engineer (Databricks) | Specialist

New
J
JobgetherData Engineering
BrazilFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Required Skills
SQLAgileApache AirflowMLFlowNosqlSparkDatabricksPySpark

Requirements

  • Demonstrated professional experience working with Databricks and modern data engineering environments.
  • Strong hands-on experience with PySpark and Apache Spark for distributed data processing.
  • Experience developing and orchestrating workflows using Apache Airflow.
  • Practical experience with Google BigQuery and cloud-based data platforms.
  • Experience integrating and using MLflow for machine learning lifecycle management.
  • Knowledge of AWS Glue and its application within data integration and processing workflows.
  • Experience working with both SQL and NoSQL databases, including technologies such as PostgreSQL, MongoDB, and Cassandra.
  • Ability to design and maintain scalable data pipelines with a strong focus on data quality, performance, automation, and reliability.
  • Experience working in Agile/Scrum environments, including sprint planning, refinement, reviews, and retrospectives.
  • Strong analytical and problem-solving abilities, with the capacity to work independently and collaboratively.
  • Knowledge of Apache Kafka for event streaming and real-time data architectures (Nice to have).
  • Experience with dbt for data transformation and analytics engineering (Nice to have).

Responsibilities

  • Build and maintain scalable, reliable data pipelines using modern distributed processing technologies to support high-quality data ingestion, transformation, and delivery.
  • Organize and manage data tables using Delta Lake and Unity Catalog, ensuring effective data management and governance.
  • Design and implement an automated, native MLOps pipeline on Databricks and Google Cloud Platform (GCP) covering the complete machine learning model lifecycle.
  • Develop processes for data preparation, feature engineering, model training, validation, registration, deployment, serving, monitoring, and automated retraining.
  • Participate in technical discovery activities, including inventorying existing machine learning models and assessing their migration requirements.
  • Develop and validate a standardized MLOps pipeline template through a pilot implementation, followed by progressive migration of models in waves based on business and technical criticality.
  • Collaborate with engineering and data teams to ensure solutions are scalable, maintainable, reliable, and aligned with technical standards.
  • Work within an agile delivery model, actively participating in sprints, refinement sessions, reviews, retrospectives, and other team rituals.
  • Continuously identify opportunities to improve data pipeline performance, automation, reliability, and operational efficiency.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now