S

Site Reliability Engineer – GPU/HPC Infrastructure (Remote – MENA)

Saturn Cloud

Remote · Full Time

Be the first to apply

Experience
Any
Salary
—
Openings
1
Posted
4 days ago
Work mode
Work from home
Resume
Required to apply

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About Saturn Cloud

Saturn Cloud delivers scalable infrastructure tailored for AI, machine learning, and data workloads. Our platform supports teams in developing, deploying, and managing compute-intensive tasks on contemporary cloud and GPU environments.

Role Overview

We seek a GPU/HPC Site Reliability Engineer based remotely within the MENA region to operate and troubleshoot our extensive GPU infrastructure, particularly supporting the Saturn Cloud Token Factory. This role emphasizes managing infrastructure around Kubernetes, focusing on GPU health, NVIDIA software, high-performance networking, topology, and distributed GPU operations impacting production inference workloads.

Key Responsibilities

  • Analyze and resolve issues affecting large-scale GPU inference infrastructure in production.
  • Diagnose problems related to NVIDIA datacenter GPUs, including drivers, CUDA compatibility, and GPU container runtimes.
  • Assess GPU health, addressing Xid errors and hardware or driver failure modes.
  • Troubleshoot PCIe configurations, NUMA settings, GPU placement, NVLink, and NVSwitch issues.
  • Manage multi-GPU and multi-node workloads effectively.
  • Investigate and resolve issues in high-performance networking and distributed communications.
  • Differentiating between application-level challenges and hardware or infrastructure faults.
  • Handle Kubernetes-based GPU workloads and NVIDIA GPU Operator/device plugins.
  • Utilize production observability tools and GPU metrics for diagnosing reliability and performance concerns.
  • Liaise with GPU-cloud and infrastructure providers on incidents requiring hardware or fabric investigations.
  • Produce clear, precise technical documentation to guide infrastructure teams during investigations.

Required Qualifications and Skills

  • Proven ability in leveraging AI coding agents and agentic development tools to enhance engineering, debugging, automation, and operations.
  • Expertise in Linux systems debugging.
  • Administration and troubleshooting experience with NVIDIA datacenter GPUs and drivers.
  • Comprehensive understanding of CUDA and corresponding driver compatibility.
  • Hands-on experience with the NVIDIA Container Toolkit/runtime.
  • Familiarity with NVML and nvidia-smi tools.
  • Strong diagnostic skills related to GPU health and error handling (including Xid errors).
  • Knowledge of PCIe topology, NUMA, GPU placement, NVLink, and NVSwitch technologies.
  • Experience with Kubernetes GPU Operator and device plugins, plus containerized GPU workload management.

High-Performance Networking Expertise

  • Experience with one or more of InfiniBand, RDMA, RoCE, NCCL, GPUDirect RDMA.
  • Understanding of NIC/GPU topology and diagnostics for NCCL and distributed workloads.
  • Ability to troubleshoot bandwidth, latency, and multi-node GPU communications.

Inference Infrastructure

While not expected to be ML researchers, candidates should understand how inference workloads utilize GPU infrastructure. Relevant experience includes working with vLLM, NVIDIA Dynamo, Triton, or similar runtimes; model loading, GPU memory management, KV cache, continuous batching; multi-GPU/multi-node inference; out-of-memory diagnosis; and performance metrics like throughput and latency.

Additional Technical Skills

  • Proficient in Kubernetes troubleshooting specific to GPU clusters.
  • Experience with Prometheus/Grafana and NVIDIA DCGM metrics.
  • Familiarity with shell scripting (Bash) and Python.
  • Experience working with containerd/Docker environments.
  • Background in managing bare-metal GPU clusters or GPU cloud infrastructure.
  • Knowledge of GPU systems such as B200/B300/H200/H100 is advantageous.

Preferred Qualifications

  • Experience with NVIDIA DCGM, Dynamo, KAI Scheduler or Grove.
  • Knowledge of Spectrum-X, Mellanox/ConnectX networking, OFED/DOCA.
  • Background in Slurm or HPC cluster administration.
  • Expertise with Kubernetes-based GPU clouds and managing large GPU fleets.
  • Familiarity with GPU burn-in, qualification, and health check tooling.

Success Criteria

Success is defined by the ability to rapidly pinpoint causes of inference workload degradation, whether software, GPU memory pressure, driver interactions, topology issues, network fabric problems, or hardware failures. Candidates will provide clear evidence to distinguish software from infrastructure faults and guide investigation efforts effectively. This role is the key point of contact when Kubernetes appears healthy but GPU workloads face issues.

Benefits and Culture

  • Fully remote role with a remote-first, high-trust work environment.
  • Direct engagement with cutting-edge GPU and AI infrastructure technologies.
  • Opportunity to tackle complex problems spanning hardware, networking, Kubernetes, and AI inference systems.
  • Competitive salary and flexible paid time off.

How they work

Communication Teamwork & Collaboration Problem Solving Attention to Detail

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer