NVIDIA

Senior Storage Production Engineer - DGX Cloud

NVIDIA

Sydney, New South Wales, Australia · Full Time

Be the first to apply

Experience
8+ yrs
Salary
Openings
1
Posted
1 week ago
Work mode
In office
Education
BSc in Computer Science or equivalent
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About the Role

This position is focused on designing, building, and maintaining high-efficiency, highly available large-scale production storage systems. It involves expertise in storage architecture, distributed storage, data management, systems engineering, and cloud technologies such as Kubernetes, containers, and virtualization. The role requires optimizing data placement and access for HPC and AI/ML workloads, ensuring systems reliability, scalability, and performance.

Key Responsibilities

  • Architect, develop, and manage expansive storage clusters with a focus on scalability, availability, and preserving data integrity.
  • Create and maintain monitoring, logging, and alert frameworks to proactively detect and address performance issues.
  • Enhance storage designs to support AI/ML workloads with low-latency access, efficient caching, and high throughput.
  • Oversee the entire lifecycle of storage services, from design through deployment, operation, and ongoing enhancement, including consultation during system builds and capacity planning.
  • Manage production storage infrastructure by monitoring availability, latency, and health using predictive analytics and AI-powered automation.
  • Boost storage performance and efficiency using compression, deduplication, data tiering strategies, and smart workload allocation.
  • Expand storage systems sustainably leveraging AI/ML-driven automation, policy-based tiering, and dynamic data migration, while ensuring security through encryption, access controls, and auditing.
  • Participate in a blameless incident response process including on-call rotation, focusing on root cause analysis and sustainable uptime.

Qualifications and Requirements

  • Bachelor’s degree or equivalent experience in Computer Science, Storage Systems, or related technical disciplines, with 8+ years of hands-on experience.
  • Proven experience with distributed and high-performance storage platforms, including clustered/parallel file systems, distributed object storage, and enterprise-grade storage.
  • In-depth knowledge of block, file, and object storage technologies, their scalability, reliability, and performance factors.
  • Familiarity with storage networking protocols like NFS, SMB, iSCSI, S3, Fibre Channel, RDMA, and NVMe over Fabrics.
  • Strong background in algorithms, data structures, software design, and automating the maintenance of large-scale Linux storage environments.
  • Proficiency in at least one of these languages for automation and tuning: C/C++, Java, Python, Go, NodeJS, or Bash.
  • Experience using infrastructure automation tools such as Ansible, Chef, Puppet, or Terraform, alongside monitoring solutions like InfluxDB, Prometheus, Grafana, and Elastic stack.
  • Excellent communication skills, strong work ethic, collaborative teamwork mindset, and dedication to quality and task completion.

Desirable Skills

  • Advanced understanding of distributed storage systems featuring replication, erasure coding, capacity planning, and high-throughput performance tuning.
  • Experience with version control, CI/CD pipelines, and infrastructure as code management.
  • Strong analytical and debugging capabilities to resolve complex storage-related problems.
  • Knowledge of networking protocols and architecture as related to storage system reliability and performance.
  • Experience in public/private cloud storage technologies including Kubernetes, OpenStack, and hybrid cloud platforms.
  • Capability to design automated storage migration, backup, and disaster recovery solutions.
  • Adaptability to diverse team environments and emerging storage technologies.

Additional Information

This role includes participation in proactive incident management and requires flexibility to join on-call rotations to uphold system reliability. It emphasizes automation, sustainable incident handling practices, and continuous infrastructure optimization. The designation references job code JR2025191.

Minimum education

Bachelor's Degree

How they work

Communication Teamwork & Collaboration Problem Solving Work Ethic

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer