Staff ML Engineer

New
J
JobgetherMachine Learning
Based in BrazilFull-TimeStaff
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Languages
Advanced English proficiency is required across conversational communication, writing, and reading.
Required Skills
AWSDockerPythonKubernetesPyTorchCI/CDLinuxTerraformMLOps

Requirements

  • Proven experience building reusable infrastructure, developer tools, or platforms that enable multiple engineers or researchers.
  • Strong proficiency in Python and Linux, including the ability to build sustainable software, automation, and services.
  • Hands-on experience with Docker, Kubernetes, CI/CD pipelines, and cloud environments such as AWS.
  • Practical understanding of the end-to-end ML lifecycle, including data preparation, experimentation, training, evaluation, model and artifact management, packaging, deployment, and monitoring.
  • Experience supporting compute-intensive or distributed workloads and diagnosing reliability, performance, resource, and cost bottlenecks.
  • Practical knowledge of modern ML frameworks such as PyTorch.
  • Ability to work with ambiguous and evolving requirements and translate them into simple, reusable engineering capabilities.
  • Strong communication and collaboration skills across research, engineering, and platform teams.
  • Advanced English proficiency is required across conversational communication, writing, and reading.
  • Strong interest in diversity, inclusion, and accessibility.

Responsibilities

  • Partner with ML researchers working on generative AI teams to identify bottlenecks and improve the speed, scalability, and reliability of research iteration.
  • Design, build, and maintain reusable research infrastructure, workflows, models, interfaces, and automation supporting experimentation, training, evaluation, data processing, and model packaging.
  • Enable reproducible experiments through consistent environments, dependency management, artifact and model versioning, configuration management, observability, and CI/CD practices.
  • Support scalable ML workloads involving large datasets, GPU clusters, distributed computing, and multiple interconnected models, services, and algorithmic components.
  • Deliver pragmatic research-enablement capabilities for immediate needs while keeping solutions aligned with the broader ML platform architecture and roadmap.
  • Act as a technical bridge between researchers and ML platform teams by translating research pain points into clear platform requirements, validating new capabilities, and supporting adoption of shared infrastructure.
  • Improve the path from research to production by making research outcomes easier to reproduce, integrate, test, and operationalize.
  • Contribute to shared ML engineering standards and architecture while promoting strong engineering practices through hands-on collaboration, technical guidance, and knowledge sharing.
  • Evaluate and introduce technologies that can materially improve research velocity, reliability, scalability, and cost efficiency.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now