Machine Learning Engineer - ML Training Platform

New
P
USA or Australia, comfortable working across timezonesFull-TimeMiddle
SalaryEquity-Heavy Package: We offer significant ownership for key technical contributors in addition to a high base salary.
Apply NowOpens the employer's application page

Job Details

Languages
Professional-level English proficiency (written and spoken)
Required Skills
AWSPythonGCPKubernetesMachine LearningAzureTerraformDistributed Systems

Requirements

  • Production experience with infrastructure-as-code (Pulumi, Terraform, CloudFormation) for multi-cloud deployments.
  • Strong proficiency with Docker, Kubernetes (EKS), and managing GPU workloads at scale.
  • Deep understanding of distributed ML training workflows: checkpointing, data sharding, model versioning, and job orchestration.
  • Experience with decentralized networking concepts: P2P, NAT traversal, and managing real bandwidth constraints.
  • Advanced Python programming skills, specifically with asyncio, concurrency, retry logic, and cloud SDKs.
  • Hands-on experience with SRE practices, observability (Prometheus/Grafana), and incident response.
  • Proven background in startup or big-tech scale service orchestration.
  • Professional-level English proficiency in written and spoken communication.
  • Alignment with the company mission regarding Protocol Learning and sovereign AI.

Responsibilities

  • Design resource management systems to provision and orchestrate multi-cloud compute across AWS, GCP, and Azure using infrastructure-as-code.
  • Architect fault-tolerant infrastructure for distributed ML, including GPU clusters, NVIDIA runtime, and S3 checkpointing.
  • Manage large-dataset streaming, model versioning, and resilient retry strategies for distributed jobs.
  • Build systems to simulate and handle real-world network conditions such as bandwidth shaping, latency injection, and packet loss.
  • Monitor node churn and maintain continuous data flow across heterogeneous worker nodes.
  • Implement SRE practices and observability using Prometheus and Grafana for system health monitoring.
View Full Description & ApplyYou'll be redirected to the employer's site
Equity-Heavy Package: We offer significant ownership for key technical contributors in addition to a high base salary.
Apply Now