Huawei Canada

Senior Researcher - Edge AI Optimization and Hardware-Aware Machine Learning

Huawei Canada

Edmonton, Alberta, Canada · Full Time

Be the first to apply

Experience
2+ yrs
Salary
Openings
1
Posted
2 weeks ago
Work mode
In office
Education
PhD
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About the Team

The Software-Hardware System Optimization Lab at Huawei Canada focuses on advancing power efficiency and performance enhancements for consumer electronics. By integrating expertise from academia and industry, the lab develops system-level optimization techniques for both software and hardware across domains like edge AI, multimedia, graphics, mobile gaming, and system software, aiming to boost user experience and competitiveness of Huawei's consumer devices.

Job Responsibilities

  • Perform research on hardware-aware neural network optimization strategies, including quantization-aware training, mixed precision, pruning, distillation, and neural architecture search.
  • Design novel methods for latency and energy-aware training goals and multi-objective optimization balancing accuracy, computational effort, and memory.
  • Create prototypes and evaluate efficient inference techniques considering device limitations such as thermal constraints, memory bandwidth, and intermittent connectivity.
  • Publish research results through internal and external channels such as papers, workshops, patents, and technical blogs.
  • Enhance inference pipelines covering preprocessing, scheduling, operator fusion, memory management, and runtime execution.
  • Collaborate in development or contribution to compiler and runtime projects (e.g., TVM, MLIR, XLA, TensorRT, ONNX Runtime, TFLite, ExecuTorch) to broaden operator support and improve performance.
  • Profile and optimize models using real device traces pinpointing bottlenecks like cache misses, memory bandwidth issues, kernel launch overhead, and CPU-NPU interactions.
  • Develop and sustain hardware-aware benchmarking and regression testing methodologies for edge devices such as ARM CPUs, mobile GPUs, DSPs, and NPUs.
  • Devise deployment strategies for heterogeneous computing platforms that involve CPUs, GPUs, and NPUs incorporating partitioning and fallback protocols.
  • Lead on-device personalization and incremental updates approaches including small adapters and efficient fine-tuning where applicable.
  • Engage closely with product engineering, platform, and hardware teams to convert device constraints into research objectives and assist in transitioning prototypes to production.
  • Mentor junior researchers and engineers by reviewing experimental designs and enforcing high standards for rigor and reproducibility.
  • Define technical roadmap initiatives including next-generation quantization, kernel optimization, model families for edge environments, and compiler advancements.

Candidate Profile and Requirements

  • PhD degree or equivalent research experience in Machine Learning, Computer Science, Electrical/Computer Engineering, or related disciplines.
  • Proficiency in Python and C/C++ (or equivalent systems languages), with experience in AI toolchains including model evaluation, orchestration, and LLM-powered tooling.
  • Hands-on experience with deep learning frameworks such as PyTorch, TensorFlow, JAX and deployment stacks like ONNX, TFLite, TensorRT, TVM, and MLIR.
  • Expertise in performance profiling including latency measurements, memory profiling, kernel bottleneck analysis, and experimental rigor.
  • Proven track record of publications at leading conferences (NeurIPS, ICML, ICLR, MLSys, ASPLOS, ISCA, MICRO) or patents related to machine learning efficiency.
  • Minimum of 2 years relevant experience from research labs or industry showing tangible impact in one or more areas: model compression (quantization, pruning, distillation), efficient architectures (MobileNet-style, MoE, efficient transformers), ML systems and compiler/runtime optimizations, or hardware-aware optimization tailored for edge deployment.
  • Experience working with edge hardware like ARM NEON, mobile GPUs, DSPs, NPUs, and microcontrollers.
  • Skills in distributed benchmarking, continuous integration for performance regression, and building reproducible experimental pipelines.
  • Understanding power and thermal constraints with measurement techniques for on-device energy consumption.
  • Experience optimizing efficient inference for large language and vision models on edge devices including techniques like KV-cache optimization, quantized attention, and speculative decoding.
  • Technical specialties including quantization (PTQ/QAT, per-channel/per-tensor, calibration, smooth quant, GPTQ methods, mixed precision), sparsity (structured and hardware-friendly pruning), compiler techniques (graph rewriting, operator lowering, scheduling, kernel autotuning), runtime methods (memory arenas, tensor lifetime analysis, shape strategies, batching), and hardware fundamentals (cache architecture, SIMD, memory bandwidth, accelerator programming).

Additional Information

Huawei Canada is dedicated to providing an equitable, inclusive, and accessible recruiting experience. Candidates requiring accommodations during any step of the hiring process should inform the team to facilitate necessary arrangements. All applications received for this role will be evaluated manually by the hiring team without using automated AI-based screening tools.

Level

Senior

Minimum education

Doctorate

🤖
Online · instant AI help
Broxer