Incident Response Engineer - Facility Operations Center
Sydney, New South Wales, Australia · Full Time
Be the first to apply
- Experience
- 5+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 2 days ago
- Work mode
- In office
- Education
- Bachelor's degree or equivalent experience
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Role
We are seeking a dedicated and experienced Incident Response Engineer to become part of our Facility Operations Center team, focusing on incident management, vendor coordination, and maintenance performance within NVIDIA’s datacenter facilities. This crucial position will help maximize uptime and minimize impacts from incidents across our data center environments.
Key Responsibilities
- Coordinate communication and manage operational incidents, maintenance activities, and reporting across NVIDIA’s datacenter sites.
- Develop and implement standards and programs to enhance reliability, including problem management and change control processes, and maintain health scores for sites with testing methods to predict and prevent failures.
- Analyze failure patterns and collaborate with AI and machine learning teams to forecast potential failures, while leading reliability studies and driving automation and process enhancements.
- Coordinate disaster recovery tests, manage audit interactions, conduct business continuity and compliance initiatives, and perform risk assessments to uphold datacenter policies and regulations.
- Lead key business metric reporting for incident response, manage tooling ownership internally and externally, and spearhead root cause analysis to improve operational documentation and prevent recurrence.
- Identify and champion process improvements and transformation projects, working with stakeholders to define problems and organize project structures.
- Foster collaboration across multidisciplinary teams, provide training and mentoring to Operations teams in using tools and systems effectively.
- Undertake additional relevant tasks and projects as assigned.
Qualifications and Experience
- Bachelor’s degree or equivalent experience in Electrical, Mechanical, Industrial, Computer, Telecommunications Engineering, Computer Science, or related business fields.
- Minimum five years of experience in data center operations or environmental, health, and safety roles.
- Expertise in reliability practices such as modeling predictions, life cycle and stress testing.
- Strong commercial and financial insight related to the cost implications of operational failures.
- Advanced numerical, statistical, and reporting skills with the ability to interpret and apply data trends effectively.
- Highly organized, goal-oriented, and capable of preparing consolidated data presentations for large audiences.
- Experienced with asset databases and Data Center Infrastructure Management (DCIM) systems to generate insights.
- Knowledge or experience managing large-scale datacenter infrastructure, including electrical, cooling, networking, or strategic roadmap development for hybrid environments.
- Advanced proficiency with Microsoft Office and Google Suite software tools.
Desirable Skills
- Experience in reliability engineering specific to electrical or mechanical cooling systems.
- Relevant certifications such as CDCMP, CMRP, CRL, CRE in maintenance and reliability fields.
- Familiarity with ISO standards related to datacenter operations and their application.
- Strong capabilities in statistical analysis, forecasting, and management information techniques.
- Proficient IT skills including advanced use of office productivity suites and adaptability to new software.
Minimum education
Bachelor's Degree
Skills
How they work
Communication
Teamwork & Collaboration
Problem Solving
Organisation