Senior Storage Engineer
R
RunpodAI Infrastructure
Remote - USAFull-TimeSenior
Salary$180,000 - $260,000
Apply NowOpens the employer's application page
Job Details
- Experience
- 8+ years in infrastructure, storage, or systems engineering
- Required Skills
- PythonKubernetesGoPrometheusRustLinux
Requirements
- 8+ years in infrastructure, storage, or systems engineering, with substantial ownership of production storage at scale.
- Deep, practical experience with at least one distributed storage system such as Ceph, MinIO, Lustre, GPFS/Spectrum Scale, MooseFS, WekaFS, VAST, or ZFS.
- Strong Linux internals and storage-stack knowledge: block layer, filesystems, NVMe, page cache, I/O schedulers, NFS/SMB, and iSCSI/NVMe-oF.
- Experience building and/or operating S3-compatible object storage services.
- Solid networking fundamentals with specific experience tuning networks for storage workloads.
- Proficiency in writing and shipping production code in Go, Python, Rust, or similar.
- Hands-on experience with observability tooling (Prometheus, Grafana, Datadog) including designing metrics.
- Track record of performance analysis and debugging under real production pressure.
- Self-starting mindset with the ability to diagnose complex performance issues independently.
- Ability to work collaboratively in a remote-first, Slack-native environment.
Responsibilities
- Own capacity, durability, availability, and performance characteristics of network volumes, local NVMe, and S3-compatible object storage.
- Tune the full I/O path: device and filesystem configuration, caching and read-ahead strategies, replication and erasure coding trade-offs, and client-side mount behavior.
- Lead capacity expansions, hardware refreshes, migrations, and rebalances without customer-visible disruption.
- Design and tune network paths for storage, including high-throughput fabric, congestion control, and NIC/offload configuration.
- Write production code (Go, Python, or similar) for storage control-plane services, provisioning workflows, and monitoring.
- Instrument the storage fleet to ensure legibility of metrics like IOPS, throughput, latency, and error rates.
- Build dashboards, SLOs, and alerts that proactively detect degradation.
View Full Description & ApplyYou'll be redirected to the employer's site