Google

Staff Software Engineer, Machine Learning Workload Optimization - TPU

Google

Singapore · Full Time

Be the first to apply

Experience
8+ yrs
Salary
Openings
1
Posted
4 hours ago
Work mode
In office
Education
Bachelor’s degree or equivalent
Eligibility
Applicants must have the current legal authorization to work in Singapore without requiring visa sponsorship.
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

Overview

Join Google’s AI and Infrastructure team in Singapore as a Staff Software Engineer dedicated to advancing machine learning workload optimization for Tensor Processing Units (TPUs). This role focuses on building scalable, high-performance ML training and inference infrastructures, supporting external customers transitioning their workloads to Google’s cloud hardware. The team offers an energetic, fast-evolving environment that encourages innovation and growth in next-gen AI infrastructure.

Key Responsibilities

  • Lead integration and optimization of stable ML stack for training and inference on cutting-edge TPU Neural Processing Interfaces (NPIs).
  • Define and architect frameworks tailored for onboarding and optimizing customer ML workloads to meet their production needs.
  • Collaborate closely with internal teams including ML research, performance engineering, and model optimization tool developers to deliver effective solutions.
  • Deliver robust, production-ready training and inference software stacks optimized for TPU platforms.
  • Develop and fine-tune large-scale reference ML models demonstrating advanced single-host and multi-host inference capabilities, particularly for extensive deployments.

Required Qualifications

  • Bachelor’s degree or equivalent experience in a relevant technical field.
  • At least 8 years of professional experience in software development.
  • A minimum of 4 years leading design and optimization of machine learning infrastructure, covering model deployment, evaluation, data processing, debugging, and fine-tuning.
  • At least 2 years working with leading-edge training frameworks like Megatron-LM, DeepSpeed and inference tools such as TensorRT-LLM, vLLM, or SGLang.

Preferred Qualifications

  • Advanced degree (Master’s or PhD) in Engineering, Computer Science, or related technical discipline.
  • Hands-on experience optimizing ML models for large-scale training and inference workloads to improve latency and throughput.
  • Familiarity with ML acceleration hardware including TPUs, GPUs, or High-Performance Computing (HPC) environments.

Additional Information

Applicants must currently have the legal right to work in Singapore without requiring visa sponsorship. The hiring process primarily involves on-site interviews.

Google is an equal opportunity employer committed to fostering diversity and inclusion. We welcome applicants irrespective of race, color, religion, sex, nationality, sexual orientation, age, disability, or veteran status. Accommodations are available upon request for applicants with disabilities.

Minimum education

Bachelor's Degree

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer