- Experience
- 4+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 2 days ago
- Work mode
- In office
- Education
- Bachelor's degree or equivalent
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About NVIDIA
For over 25 years, NVIDIA has reshaped computer graphics, PC gaming, and accelerated computing through groundbreaking technology and exceptional talent. We are now spearheading the AI revolution to redefine computing where GPUs power devices like computers, robots, and autonomous vehicles to perceive and interact with the world. At NVIDIA, you'll be part of an inclusive, diverse environment inspiring you to achieve your best work and make a global impact.
Role Overview
We are seeking a Senior AI/HPC Engineer to join our Infrastructure Specialist team. This position involves collaborating with academic and commercial clients across the globe who use NVIDIA products to revolutionize deep learning, data analytics, and data centers by building some of the largest and most powerful AI/HPC systems worldwide. The role demands excellent interpersonal skills as you will partner with customers, partners, and internal teams to analyze, specify, and execute large-scale AI/HPC infrastructure projects involving networking, system design, and automation while serving as the customer-facing expert.
Key Responsibilities
- Deploy, administer, and maintain AI/HPC infrastructures on Linux platforms for both new and current customers.
- Serve as a subject matter expert during customer planning discussions, guiding projects through to completion.
- Prepare comprehensive documentation and facilitate knowledge transfers to support customer adoption of sophisticated systems.
- Collaborate with internal teams by reporting bugs, documenting workarounds, and recommending improvements.
Qualifications and Requirements
- Bachelor’s degree in Computer Science, Electrical Engineering, or related discipline, or equivalent professional experience.
- Minimum of four years offering advanced support and deployment expertise for hardware and software solutions.
- Strong familiarity with AI Factory/HPC ecosystems, including Linux system administration, process and kernel management, boot procedures, troubleshooting, performance monitoring, and optimization.
- Experience managing compute clusters and utilizing scheduling systems such as SLURM, LSF, or PBS.
- Proficient scripting capabilities.
- Excellent communication skills in English, both written and verbal, coupled with strong interpersonal abilities to resolve customer challenges quickly.
- Highly organized with strong multitasking and prioritization skills, functioning well independently.
Preferred Skills
- Linux industry certification credentials.
- Hands-on experience with InfiniBand and Ethernet networking technologies.
- Expertise in GPU-centric hardware and software ecosystems.
- Experience with Message Passing Interface (MPI) frameworks.
- Background in automation tools such as Ansible, Salt, or Puppet.
Additional Information
NVIDIA is recognized globally as a leading employer known for innovations in Artificial Intelligence, High-Performance Computing, and Visualization. We provide competitive compensation, comprehensive benefits, and a workplace culture that embraces diversity, inclusion, and work flexibility. We invite passionate and autonomous professionals ready to tackle challenges and excel in a dynamic software design team to join us. NVIDIA is an equal opportunity employer dedicated to creating a supportive and empowering environment for everyone.
Level
Senior
Minimum education
Bachelor's Degree