Senior Storage Production Engineer - DGX Cloud
New
J
JobgetherCloud Infrastructure
Based in AustraliaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 8+ years
- Required Skills
- PythonKubernetesC++GoGrafanaPrometheusLinuxTerraformAnsible
Requirements
- Bachelor’s degree or equivalent practical experience in Computer Science, Storage Systems, or a related technical discipline.
- 8+ years of practical experience working with production infrastructure, storage systems, or closely related technologies.
- Strong experience with distributed and high-performance storage solutions, including clustered and parallel file systems, distributed object storage, and enterprise storage platforms.
- Deep understanding of block, file, and object storage technologies, including their scalability, reliability, performance characteristics, and operational requirements.
- Experience with storage networking protocols such as NFS, SMB, iSCSI, S3, Fibre Channel, RDMA, and NVMe over Fabrics.
- Strong knowledge of algorithms, data structures, complexity analysis, software design, and automation of large-scale Linux-based storage environments.
- Professional programming experience in one or more relevant languages, such as C/C++, Java, Python, Go, NodeJS, or Bash.
- Hands-on experience with infrastructure configuration and automation tools such as Ansible, Chef, Puppet, or Terraform.
- Experience with observability and monitoring technologies such as InfluxDB, Prometheus, Grafana, or the Elastic stack.
- Strong debugging and systematic problem-solving abilities for complex storage and infrastructure failures.
- Strong written and verbal communication skills.
Responsibilities
- Design, implement, operate, and continuously improve large-scale storage clusters with strong scalability, availability, reliability, and data integrity.
- Develop and maintain monitoring, logging, alerting, and observability systems to proactively identify storage health and performance issues.
- Optimize storage architectures for AI/ML workloads, focusing on low-latency data access, efficient caching, high throughput, and predictable performance.
- Manage the complete lifecycle of storage services, from architecture and initial design through deployment, production operation, and ongoing optimization.
- Automate storage operations and implement intelligent mechanisms for fault detection, remediation, data placement, and system optimization.
- Participate in incident response, sustainable operational practices, and blameless root-cause analysis.
- Participate in an on-call rotation supporting critical storage and production systems.
View Full Description & ApplyYou'll be redirected to the employer's site