Machine Learning Engineer - ML Training Platform
New
P
Pluralis ResearchAI Research
USA or Australia, comfortable working across timezonesFull-TimeMiddle
SalaryEquity-Heavy Package: We offer significant ownership for key technical contributors in addition to a high base salary.
Apply NowOpens the employer's application page
Job Details
- Languages
- Professional-level English proficiency (written and spoken)
- Required Skills
- AWSPythonGCPKubernetesMachine LearningAzureTerraformDistributed Systems
Requirements
- Production experience with infrastructure-as-code (Pulumi, Terraform, CloudFormation) for multi-cloud deployments.
- Strong proficiency with Docker, Kubernetes (EKS), and managing GPU workloads at scale.
- Deep understanding of distributed ML training workflows: checkpointing, data sharding, model versioning, and job orchestration.
- Experience with decentralized networking concepts: P2P, NAT traversal, and managing real bandwidth constraints.
- Advanced Python programming skills, specifically with asyncio, concurrency, retry logic, and cloud SDKs.
- Hands-on experience with SRE practices, observability (Prometheus/Grafana), and incident response.
- Proven background in startup or big-tech scale service orchestration.
- Professional-level English proficiency in written and spoken communication.
- Alignment with the company mission regarding Protocol Learning and sovereign AI.
Responsibilities
- Design resource management systems to provision and orchestrate multi-cloud compute across AWS, GCP, and Azure using infrastructure-as-code.
- Architect fault-tolerant infrastructure for distributed ML, including GPU clusters, NVIDIA runtime, and S3 checkpointing.
- Manage large-dataset streaming, model versioning, and resilient retry strategies for distributed jobs.
- Build systems to simulate and handle real-world network conditions such as bandwidth shaping, latency injection, and packet loss.
- Monitor node churn and maintain continuous data flow across heterogeneous worker nodes.
- Implement SRE practices and observability using Prometheus and Grafana for system health monitoring.
View Full Description & ApplyYou'll be redirected to the employer's site