Site Reliability Engineer – GPU/HPC Infrastructure (Remote – MENA)
Remote · Full Time
Be the first to apply
- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 4 days ago
- Work mode
- Work from home
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Saturn Cloud
Saturn Cloud delivers scalable infrastructure tailored for AI, machine learning, and data workloads. Our platform supports teams in developing, deploying, and managing compute-intensive tasks on contemporary cloud and GPU environments.
Role Overview
We seek a GPU/HPC Site Reliability Engineer based remotely within the MENA region to operate and troubleshoot our extensive GPU infrastructure, particularly supporting the Saturn Cloud Token Factory. This role emphasizes managing infrastructure around Kubernetes, focusing on GPU health, NVIDIA software, high-performance networking, topology, and distributed GPU operations impacting production inference workloads.
Key Responsibilities
- Analyze and resolve issues affecting large-scale GPU inference infrastructure in production.
- Diagnose problems related to NVIDIA datacenter GPUs, including drivers, CUDA compatibility, and GPU container runtimes.
- Assess GPU health, addressing Xid errors and hardware or driver failure modes.
- Troubleshoot PCIe configurations, NUMA settings, GPU placement, NVLink, and NVSwitch issues.
- Manage multi-GPU and multi-node workloads effectively.
- Investigate and resolve issues in high-performance networking and distributed communications.
- Differentiating between application-level challenges and hardware or infrastructure faults.
- Handle Kubernetes-based GPU workloads and NVIDIA GPU Operator/device plugins.
- Utilize production observability tools and GPU metrics for diagnosing reliability and performance concerns.
- Liaise with GPU-cloud and infrastructure providers on incidents requiring hardware or fabric investigations.
- Produce clear, precise technical documentation to guide infrastructure teams during investigations.
Required Qualifications and Skills
- Proven ability in leveraging AI coding agents and agentic development tools to enhance engineering, debugging, automation, and operations.
- Expertise in Linux systems debugging.
- Administration and troubleshooting experience with NVIDIA datacenter GPUs and drivers.
- Comprehensive understanding of CUDA and corresponding driver compatibility.
- Hands-on experience with the NVIDIA Container Toolkit/runtime.
- Familiarity with NVML and nvidia-smi tools.
- Strong diagnostic skills related to GPU health and error handling (including Xid errors).
- Knowledge of PCIe topology, NUMA, GPU placement, NVLink, and NVSwitch technologies.
- Experience with Kubernetes GPU Operator and device plugins, plus containerized GPU workload management.
High-Performance Networking Expertise
- Experience with one or more of InfiniBand, RDMA, RoCE, NCCL, GPUDirect RDMA.
- Understanding of NIC/GPU topology and diagnostics for NCCL and distributed workloads.
- Ability to troubleshoot bandwidth, latency, and multi-node GPU communications.
Inference Infrastructure
While not expected to be ML researchers, candidates should understand how inference workloads utilize GPU infrastructure. Relevant experience includes working with vLLM, NVIDIA Dynamo, Triton, or similar runtimes; model loading, GPU memory management, KV cache, continuous batching; multi-GPU/multi-node inference; out-of-memory diagnosis; and performance metrics like throughput and latency.
Additional Technical Skills
- Proficient in Kubernetes troubleshooting specific to GPU clusters.
- Experience with Prometheus/Grafana and NVIDIA DCGM metrics.
- Familiarity with shell scripting (Bash) and Python.
- Experience working with containerd/Docker environments.
- Background in managing bare-metal GPU clusters or GPU cloud infrastructure.
- Knowledge of GPU systems such as B200/B300/H200/H100 is advantageous.
Preferred Qualifications
- Experience with NVIDIA DCGM, Dynamo, KAI Scheduler or Grove.
- Knowledge of Spectrum-X, Mellanox/ConnectX networking, OFED/DOCA.
- Background in Slurm or HPC cluster administration.
- Expertise with Kubernetes-based GPU clouds and managing large GPU fleets.
- Familiarity with GPU burn-in, qualification, and health check tooling.
Success Criteria
Success is defined by the ability to rapidly pinpoint causes of inference workload degradation, whether software, GPU memory pressure, driver interactions, topology issues, network fabric problems, or hardware failures. Candidates will provide clear evidence to distinguish software from infrastructure faults and guide investigation efforts effectively. This role is the key point of contact when Kubernetes appears healthy but GPU workloads face issues.
Benefits and Culture
- Fully remote role with a remote-first, high-trust work environment.
- Direct engagement with cutting-edge GPU and AI infrastructure technologies.
- Opportunity to tackle complex problems spanning hardware, networking, Kubernetes, and AI inference systems.
- Competitive salary and flexible paid time off.