Staff ML Engineer
New
J
JobgetherMachine Learning
Based in BrazilFull-TimeStaff
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Languages
- Advanced English proficiency is required across conversational communication, writing, and reading.
- Required Skills
- AWSDockerPythonKubernetesPyTorchCI/CDLinuxTerraformMLOps
Requirements
- Proven experience building reusable infrastructure, developer tools, or platforms that enable multiple engineers or researchers.
- Strong proficiency in Python and Linux, including the ability to build sustainable software, automation, and services.
- Hands-on experience with Docker, Kubernetes, CI/CD pipelines, and cloud environments such as AWS.
- Practical understanding of the end-to-end ML lifecycle, including data preparation, experimentation, training, evaluation, model and artifact management, packaging, deployment, and monitoring.
- Experience supporting compute-intensive or distributed workloads and diagnosing reliability, performance, resource, and cost bottlenecks.
- Practical knowledge of modern ML frameworks such as PyTorch.
- Ability to work with ambiguous and evolving requirements and translate them into simple, reusable engineering capabilities.
- Strong communication and collaboration skills across research, engineering, and platform teams.
- Advanced English proficiency is required across conversational communication, writing, and reading.
- Strong interest in diversity, inclusion, and accessibility.
Responsibilities
- Partner with ML researchers working on generative AI teams to identify bottlenecks and improve the speed, scalability, and reliability of research iteration.
- Design, build, and maintain reusable research infrastructure, workflows, models, interfaces, and automation supporting experimentation, training, evaluation, data processing, and model packaging.
- Enable reproducible experiments through consistent environments, dependency management, artifact and model versioning, configuration management, observability, and CI/CD practices.
- Support scalable ML workloads involving large datasets, GPU clusters, distributed computing, and multiple interconnected models, services, and algorithmic components.
- Deliver pragmatic research-enablement capabilities for immediate needs while keeping solutions aligned with the broader ML platform architecture and roadmap.
- Act as a technical bridge between researchers and ML platform teams by translating research pain points into clear platform requirements, validating new capabilities, and supporting adoption of shared infrastructure.
- Improve the path from research to production by making research outcomes easier to reproduce, integrate, test, and operationalize.
- Contribute to shared ML engineering standards and architecture while promoting strong engineering practices through hands-on collaboration, technical guidance, and knowledge sharing.
- Evaluate and introduce technologies that can materially improve research velocity, reliability, scalability, and cost efficiency.
View Full Description & ApplyYou'll be redirected to the employer's site