Engineering Manager, Production Platform and Orchestration

New
J
JobgetherAI infrastructure
Position is remote within the United StatesFull-TimeManager
SalaryBase salary range of $200,000–$300,000. Competitive equity as part of the overall compensation package.
Apply NowOpens the employer's application page

Job Details

Required Skills
LinuxDistributed SystemsNetworking

Requirements

  • Demonstrated success managing and growing engineering teams responsible for distributed systems, cloud infrastructure, production platforms, SRE, or closely related technical domains.
  • Strong technical judgment across Linux systems, networking, orchestration, deployment systems, observability, and production reliability.
  • Ability to turn ambiguous operational challenges into clear roadmaps, ownership models, and measurable engineering outcomes.
  • Experience recruiting, coaching, developing, and retaining engineers across experience levels.
  • Ability to work across software, hardware, data center, customer, and business boundaries.
  • Experience operating services with meaningful availability expectations, including on-call programs, incident management, root-cause analysis, and reliability planning.
  • Hands-on leadership style with technical and operational depth to engage in architecture and production discussions.
  • Experience with control planes, schedulers, placement systems, capacity-management systems, or multi-rack orchestration is highly valuable.
  • Experience with GPU, FPGA, ASIC, or other accelerator fleets in production is strongly preferred.
  • Background in data center deployment, hardware lifecycle management, firmware coordination, sparing/RMA processes, or production networking is advantageous.
  • Demonstrated ability to scale infrastructure from early deployments to multiple sites or hundreds of systems.
  • Proven automation experience that has reduced operational toil, incident frequency, recovery time, or manual intervention.

Responsibilities

  • Lead, coach, and grow a distributed engineering team focused on production platforms, fleet orchestration, reliability, and operational automation.
  • Establish ownership boundaries, develop technical leaders, recruit engineers, and scale the organization with the infrastructure footprint.
  • Translate customer and business priorities into an engineering roadmap.
  • Define technical strategy for provisioning, deployment, environment lifecycle, fleet health, capacity, upgrades, rollback, and production control-plane capabilities.
  • Build orchestration systems for inventory, placement, configuration, health management, and multi-system operations across a heterogeneous accelerator environment.
  • Establish service-level objectives, operational metrics, alerting standards, and sustainable on-call practices.
  • Own incident, escalation, postmortem, corrective-action, launch-readiness, and reliability-review processes.
  • Drive automation for deployment, upgrades, remediation, diagnostics, capacity planning, and support workflows.
  • Partner with hardware and data center teams on rack bring-up, networking, firmware, sparing, failure handling, and infrastructure transitions.
  • Monitor fleet health, incident patterns, deployment reliability, operational toil, and critical failure points to prioritize platform investments.
View Full Description & ApplyYou'll be redirected to the employer's site
Base salary range of $200,000–$300,000. Competitive equity as part of the overall compensation package.
Apply Now