This company operates with a remote-first culture, allowing team members to work from anywhere. Team members are distributed across the U.S. and Canada.
Partner with Data Center Management and Host/Supply teams to maintain capacity and utilization review cadences.
Own escalation paths with infrastructure partners to resolve capacity, performance, or contractual issues.
Translate capacity and utilization data into actionable tradeoffs regarding supply, consolidation, and cost.
Develop and track project plans across Engineering, Product, and Supply stakeholders to manage dependencies and timelines.
Act as the primary point of contact to identify risks, remove roadblocks, and maintain stakeholder alignment.
Evaluate infrastructure tools and vendors to ensure they meet scalability and performance constraints.
Execute post-program reviews to capture lessons learned and improve future cycles.
SaaS
Showing 1 of 21 positions
About Runpod
Runpod is the AI Developer Cloud, enabling over one million developers to build, deploy, and scale AI models in seconds. You will join a company processing 20 billion inference requests, transforming AI development for indie researchers to enterprise teams. Runpod offers cost-effective GPU infrastructure, simplifying the complex world of AI deployment.
How We Work
You will join a small, remote-first team distributed across the U.S., Canada, Europe, and India. We prioritize ownership and move quickly. We ship work relied on by over a million developers daily. Our culture fosters collaboration and supports transparent decision-making. We value empowering leadership, encouraging every team member to contribute ideas.
Engineering at Runpod
At Runpod, you will solve critical challenges in GPU cloud computing and AI infrastructure. Engineers build high-performance systems for large-scale AI workloads, from optimizing low-latency HPC networks using InfiniBand and RoCE to architecting distributed storage with NVMe-oF. You will work with Python, Typescript, Go Lang, Docker, and Kubernetes. You can contribute to security at the kernel level, defending against tenant breakouts and privilege escalation. Our technical approach ensures our platform remains fast, reliable, and massively scalable for AI innovation.
Why Join Us
Impact over one million developers and 20 billion inference requests daily.
Receive meaningful equity in a fast-growing company, sharing in our success.
Work in a remote-first environment that emphasizes ownership and rapid iteration.
Build infrastructure powering the next generation of AI systems.
Contribute to a culture that values urgency, deep care, and significant impact at scale.
Benefits & Perks
Meaningful equity (stock options) for all team members.
Generous medical, dental & vision plans (100% coverage for employees, partial for dependents).
Flexible PTO to recharge when you need it.
$1,200 Home Office & Equipment Stipend to set up your ideal workspace.