B
GBU/TBU Model Optimization AI Inference Engineer
Singapore · Full Time
Be the first to apply
- Experience
- 3+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 6 days ago
- Work mode
- In office
- Education
- Bachelor's degree
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Job Overview
This role involves deploying and optimizing deep learning models across heterogeneous hardware platforms, specifically NVIDIA GPUs and TPUs, focusing on improving inference performance and enabling continuous enhancements.
Key Responsibilities
- Manage cross-platform model deployment for deep learning models on GPU (NVIDIA) and TPU hardware focusing on inference, performance tuning, and iterative improvements.
- Deeply understand hardware characteristics of GPUs (CUDA/core architecture) and TPUs (systolic arrays/MXU) to identify and resolve performance bottlenecks like memory access, operator efficiency, and chip utilization, enhancing throughput while reducing latency.
- Design and implement efficient model inference solutions including model format conversion, quantization (INT8/FP16/BF16), operator fusion, and graph optimization; perform comparative analysis of inference cost-effectiveness and applicability between GPU and TPU platforms.
- Collaborate closely with algorithm research teams to co-design model architectures and inference performance, balancing algorithm accuracy and cross-platform efficiency.
- Develop and refine best practices for GPU/TPU inference deployment and explore building unified inference deployment toolchains or abstraction layers to simplify adaptation for business applications.
Candidate Qualifications
- Bachelor’s degree or higher in Computer Science, Electronic Engineering, Automation, or related fields.
- More than 3 years of experience in AI infrastructure, inference engines, or high-performance computing.
- Solid understanding of GPU or TPU architectures; proficiency with NVIDIA GPU CUDA programming model and low-level mechanisms; in-depth knowledge of TPU or other AI accelerators’ working principles.
- Proficiency in at least one mainstream deep learning framework such as PyTorch, TensorFlow, or JAX, with familiarity of the complete pipeline from training to deployment.
- Strong programming skills in C++ and Python, with a focus on system programming and performance tuning.
- Familiarity with GPU ecosystem tools like TensorRT, cuDNN, NCCL, or Triton Inference Server; or TPU tools like XLA, PyTorch/XLA, or JAX compiler and runtime environments.
Minimum education
Bachelor's Degree