Engineering Manager, Production Platform and Orchestration
New
J
JobgetherAI infrastructure
Position is remote within the United StatesFull-TimeManager
SalaryBase salary range of $200,000–$300,000. Competitive equity as part of the overall compensation package.
Apply NowOpens the employer's application page
Job Details
- Required Skills
- LinuxDistributed SystemsNetworking
Requirements
- Demonstrated success managing and growing engineering teams responsible for distributed systems, cloud infrastructure, production platforms, SRE, or closely related technical domains.
- Strong technical judgment across Linux systems, networking, orchestration, deployment systems, observability, and production reliability.
- Ability to turn ambiguous operational challenges into clear roadmaps, ownership models, and measurable engineering outcomes.
- Experience recruiting, coaching, developing, and retaining engineers across experience levels.
- Ability to work across software, hardware, data center, customer, and business boundaries.
- Experience operating services with meaningful availability expectations, including on-call programs, incident management, root-cause analysis, and reliability planning.
- Hands-on leadership style with technical and operational depth to engage in architecture and production discussions.
- Experience with control planes, schedulers, placement systems, capacity-management systems, or multi-rack orchestration is highly valuable.
- Experience with GPU, FPGA, ASIC, or other accelerator fleets in production is strongly preferred.
- Background in data center deployment, hardware lifecycle management, firmware coordination, sparing/RMA processes, or production networking is advantageous.
- Demonstrated ability to scale infrastructure from early deployments to multiple sites or hundreds of systems.
- Proven automation experience that has reduced operational toil, incident frequency, recovery time, or manual intervention.
Responsibilities
- Lead, coach, and grow a distributed engineering team focused on production platforms, fleet orchestration, reliability, and operational automation.
- Establish ownership boundaries, develop technical leaders, recruit engineers, and scale the organization with the infrastructure footprint.
- Translate customer and business priorities into an engineering roadmap.
- Define technical strategy for provisioning, deployment, environment lifecycle, fleet health, capacity, upgrades, rollback, and production control-plane capabilities.
- Build orchestration systems for inventory, placement, configuration, health management, and multi-system operations across a heterogeneous accelerator environment.
- Establish service-level objectives, operational metrics, alerting standards, and sustainable on-call practices.
- Own incident, escalation, postmortem, corrective-action, launch-readiness, and reliability-review processes.
- Drive automation for deployment, upgrades, remediation, diagnostics, capacity planning, and support workflows.
- Partner with hardware and data center teams on rack bring-up, networking, firmware, sparing, failure handling, and infrastructure transitions.
- Monitor fleet health, incident patterns, deployment reliability, operational toil, and critical failure points to prioritize platform investments.
View Full Description & ApplyYou'll be redirected to the employer's site