Senior Storage Production Engineer - DGX Cloud
Sydney, New South Wales, Australia · Full Time
Be the first to apply
- Experience
- 8+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 1 week ago
- Work mode
- In office
- Education
- BSc in Computer Science or equivalent
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Role
This position is focused on designing, building, and maintaining high-efficiency, highly available large-scale production storage systems. It involves expertise in storage architecture, distributed storage, data management, systems engineering, and cloud technologies such as Kubernetes, containers, and virtualization. The role requires optimizing data placement and access for HPC and AI/ML workloads, ensuring systems reliability, scalability, and performance.
Key Responsibilities
- Architect, develop, and manage expansive storage clusters with a focus on scalability, availability, and preserving data integrity.
- Create and maintain monitoring, logging, and alert frameworks to proactively detect and address performance issues.
- Enhance storage designs to support AI/ML workloads with low-latency access, efficient caching, and high throughput.
- Oversee the entire lifecycle of storage services, from design through deployment, operation, and ongoing enhancement, including consultation during system builds and capacity planning.
- Manage production storage infrastructure by monitoring availability, latency, and health using predictive analytics and AI-powered automation.
- Boost storage performance and efficiency using compression, deduplication, data tiering strategies, and smart workload allocation.
- Expand storage systems sustainably leveraging AI/ML-driven automation, policy-based tiering, and dynamic data migration, while ensuring security through encryption, access controls, and auditing.
- Participate in a blameless incident response process including on-call rotation, focusing on root cause analysis and sustainable uptime.
Qualifications and Requirements
- Bachelor’s degree or equivalent experience in Computer Science, Storage Systems, or related technical disciplines, with 8+ years of hands-on experience.
- Proven experience with distributed and high-performance storage platforms, including clustered/parallel file systems, distributed object storage, and enterprise-grade storage.
- In-depth knowledge of block, file, and object storage technologies, their scalability, reliability, and performance factors.
- Familiarity with storage networking protocols like NFS, SMB, iSCSI, S3, Fibre Channel, RDMA, and NVMe over Fabrics.
- Strong background in algorithms, data structures, software design, and automating the maintenance of large-scale Linux storage environments.
- Proficiency in at least one of these languages for automation and tuning: C/C++, Java, Python, Go, NodeJS, or Bash.
- Experience using infrastructure automation tools such as Ansible, Chef, Puppet, or Terraform, alongside monitoring solutions like InfluxDB, Prometheus, Grafana, and Elastic stack.
- Excellent communication skills, strong work ethic, collaborative teamwork mindset, and dedication to quality and task completion.
Desirable Skills
- Advanced understanding of distributed storage systems featuring replication, erasure coding, capacity planning, and high-throughput performance tuning.
- Experience with version control, CI/CD pipelines, and infrastructure as code management.
- Strong analytical and debugging capabilities to resolve complex storage-related problems.
- Knowledge of networking protocols and architecture as related to storage system reliability and performance.
- Experience in public/private cloud storage technologies including Kubernetes, OpenStack, and hybrid cloud platforms.
- Capability to design automated storage migration, backup, and disaster recovery solutions.
- Adaptability to diverse team environments and emerging storage technologies.
Additional Information
This role includes participation in proactive incident management and requires flexibility to join on-call rotations to uphold system reliability. It emphasizes automation, sustainable incident handling practices, and continuous infrastructure optimization. The designation references job code JR2025191.
Minimum education
Bachelor's Degree