Singtel

DevOps Engineer - GPU-as-a-Service (GPUaaS)

Singtel

Singapore · Full Time

Be the first to apply

Experience
Any
Salary
Openings
1
Posted
1 day ago
Work mode
In office
Education
Bachelor's degree in Computer Science or related field
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About Singtel Digital InfraCo RE:AI

Singtel Digital InfraCo’s RE:AI division is focused on creating Asia’s most advanced and sustainable AI infrastructure ecosystem. This initiative supports enterprises, research bodies, and digital businesses by accelerating innovation through responsible, high-performance AI computing and connectivity services.

Role Overview

Join as a DevOps Engineer contributing to Singtel’s GPU-as-a-Service platform, facilitating AI and High Performance Computing (HPC) capabilities for clients. The role blends practical physical data center work with software solutions, aimed at advancing GPU infrastructure for AI and HPC workloads. Ideal candidates are proactive, adaptable to evolving environments, and eager to enhance GPU cloud platforms through continuous improvements.

Key Responsibilities

  • Architect, deploy, and maintain extensive distributed GPU clusters dedicated to AI and machine learning tasks.
  • Automate and manage the allocation of GPU resources across on-premises and cloud infrastructures.
  • Develop, implement, and oversee CI/CD pipelines supporting AI models and GPU-accelerated applications.
  • Continuously monitor the health, performance, usage, and uptime of GPU clusters.
  • Enhance provisioning and management processes through automation technologies.
  • Diagnose and resolve system-level issues involving Slurm, Kubernetes, GPU drivers, CUDA, and InfiniBand networking.
  • Optimize operating system, driver, network, and library configurations to maximize AI workload performance.
  • Conduct benchmarking assessments of GPU clusters while staying current with advancements in GPU technologies.
  • Implement monitoring and logging solutions using tools like Zabbix, Prometheus, and NVIDIA DCGM.
  • Apply security best practices in managing a multi-tenant GPU-as-a-Service environment.
  • Collaborate across teams to streamline workflows, enhance cooperation, and provide expert guidance to GPU system users.
  • Partner with senior engineers to identify process bottlenecks and refine operational procedures for AI/HPC GPU cloud services.
  • Engage in rotational or scheduled shifts to support platform operation needs.

Required Qualifications and Skills

  • Bachelor’s degree in Computer Science, Information Technology, Systems Engineering, or related disciplines.
  • Proficient with Linux system administration (Ubuntu, CentOS, Rocky Linux, etc.).
  • Experience utilizing DevOps tools including Jenkins, Kubernetes, Ansible, and Terraform.
  • Strong grasp of CI/CD pipelines, automation strategies, and monitoring techniques.
  • Competent in scripting languages such as Python and Bash.
  • Implemented monitoring frameworks like Zabbix and Prometheus.
  • Knowledge of AI frameworks including TensorFlow and PyTorch.
  • Understanding of cloud service models (IaaS, PaaS), GPU hardware architecture, specifically NVIDIA GPUs.
  • Excellent communication and presentation skills in English.
  • Collaborative team player with cross-functional coordination experience.
  • Analytical problem-solving capabilities focused on system optimization.

Preferred Additional Expertise

  • Insight into collective communication protocols such as MPI, RDMA, NCCL related to GPU acceleration.
  • Familiarity with DevOps and MLOps tools for GPU clusters including Docker containers, Kubernetes, and data center deployment.
  • Experience with HPC workload schedulers, notably Slurm.
  • Understanding of AI & HPC networking technologies including InfiniBand, RoCE, and DPUs.
  • Hands-on system experience with NVIDIA GPU and associated software development kits.
  • Comprehension of interaction between AI/HPC workloads and GPU hardware/software infrastructures.

Additional Information

The position requires engagement in shift work on a rotational or scheduled basis to maintain service operations.

Minimum education

Bachelor's Degree

Tools & software

Kubernetes required PyTorch required TensorFlow required Zabbix required

How they work

Communication Teamwork & Collaboration Problem Solving Adaptability

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer