Data Engineer (Databricks) | Specialist
New
J
JobgetherData Engineering
BrazilFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- SQLAgileApache AirflowMLFlowNosqlSparkDatabricksPySpark
Requirements
- Demonstrated professional experience working with Databricks and modern data engineering environments.
- Strong hands-on experience with PySpark and Apache Spark for distributed data processing.
- Experience developing and orchestrating workflows using Apache Airflow.
- Practical experience with Google BigQuery and cloud-based data platforms.
- Experience integrating and using MLflow for machine learning lifecycle management.
- Knowledge of AWS Glue and its application within data integration and processing workflows.
- Experience working with both SQL and NoSQL databases, including technologies such as PostgreSQL, MongoDB, and Cassandra.
- Ability to design and maintain scalable data pipelines with a strong focus on data quality, performance, automation, and reliability.
- Experience working in Agile/Scrum environments, including sprint planning, refinement, reviews, and retrospectives.
- Strong analytical and problem-solving abilities, with the capacity to work independently and collaboratively.
- Knowledge of Apache Kafka for event streaming and real-time data architectures (Nice to have).
- Experience with dbt for data transformation and analytics engineering (Nice to have).
Responsibilities
- Build and maintain scalable, reliable data pipelines using modern distributed processing technologies to support high-quality data ingestion, transformation, and delivery.
- Organize and manage data tables using Delta Lake and Unity Catalog, ensuring effective data management and governance.
- Design and implement an automated, native MLOps pipeline on Databricks and Google Cloud Platform (GCP) covering the complete machine learning model lifecycle.
- Develop processes for data preparation, feature engineering, model training, validation, registration, deployment, serving, monitoring, and automated retraining.
- Participate in technical discovery activities, including inventorying existing machine learning models and assessing their migration requirements.
- Develop and validate a standardized MLOps pipeline template through a pilot implementation, followed by progressive migration of models in waves based on business and technical criticality.
- Collaborate with engineering and data teams to ensure solutions are scalable, maintainable, reliable, and aligned with technical standards.
- Work within an agile delivery model, actively participating in sprints, refinement sessions, reviews, retrospectives, and other team rituals.
- Continuously identify opportunities to improve data pipeline performance, automation, reliability, and operational efficiency.
View Full Description & ApplyYou'll be redirected to the employer's site