Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start
Singapore · Full Time
Be the first to apply
- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 3 hours ago
- Work mode
- In office
- Education
- Bachelor's or Master's in computing or related discipline
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Team
Data AML is ByteDance's mid-platform for Machine Learning, supporting training and inference systems across multiple business lines, including Douyin, Jinri Toutiao, and Xigua Video. It delivers robust ML computational capabilities and pioneers innovative algorithms to solve key business challenges.
Position Overview
We seek talented graduates joining in 2027 to engage in ambitious projects involving complex problem solving and career growth within ByteDance’s inspiring environment. Candidates must commit to onboarding by the end of 2027 and clearly indicate their availability and graduation date in their resume.
Application Details
Applicants may apply up to two positions within ByteDance and its affiliates worldwide, with applications reviewed continuously; early application is recommended.
Key Responsibilities
- Revamp foundational architecture of the large model inference engine and optimize GPU performance end-to-end via operator fusion, compilation enhancements, and deep optimization of GPU memory access and compute pipelines to remove bottlenecks, increase single-GPU throughput, and reduce latency.
- Ensure compatibility across diverse GPU/NPU architectures, enhancing the inference engine’s adaptability and building a high-performance, minimal-loss base for large-scale model inference.
- Lead design, development, and enhancement of distributed parallel strategies including tensor, pipeline, sequence, and expert parallelisms to manage multi-GPU model deployment challenges such as communication overhead, load imbalance, and efficiency.
- Stay current with cutting-edge technologies in global model inference, GPU high-performance computing, distributed parallelism, and caching; benchmark with industry-standard frameworks like vLLM and TensorRT-LLM to innovate and improve inference system performance and cost-effectiveness.
Minimum Qualifications
- Applicants completing or recently having completed a Bachelor's or Master's degree in computing or related fields.
- Strong fundamentals in computer low-level systems, proficient in C/C++, Python, CUDA programming, and knowledgeable about GPU hardware architecture and memory models.
- Expertise in developing and optimizing core deep learning operators (matrix operations, normalization, activation functions) including operator reconstruction, memory optimization, vectorization, and precision alignment to ensure reliability and performance in inference.
- Understanding of deep learning inference compilation including graph optimization, operator fusion, constant folding, memory reuse, scheduling, and quantization to streamline inference, reduce GPU memory usage and latency, and boost throughput efficiency.
- Experience with GPU profiling and analysis tools such as Nsight and Profiler to identify bottlenecks and design systematic software-hardware optimizations suitable for high-concurrency, low-latency industrial inference applications.
- Effective collaboration, communication, presentation, and technical documentation skills, combined with strong responsibility, resilience, and capacity to resolve complex technical challenges and drive project execution.
Preferred Qualifications
- Comprehensive grasp of large model inference principles and model parallelism technologies, including experience with distributed inference solutions managing multi-GPU communication, load balancing, and parallel efficiency.
- Experience contributing to the development or performance tuning of mainstream large model inference frameworks such as vLLM, SGLang, or TensorRT-LLM is advantageous.
About ByteDance
Founded in 2012, ByteDance is committed to inspiring creativity and enriching life. Our products—including TikTok, Lemon8, CapCut, Pico, and China-specific platforms like Toutiao and Douyin—facilitate connection, content consumption, and creation.
Why Join Us?
Creativity drives our mission. Our diverse global teams foster authentic self-expression, discovery, and community building. We prioritize curiosity, humility, and impactful innovation, maintaining an "Always Day 1" mindset to achieve breakthroughs together.
Diversity & Inclusion
ByteDance values diverse skills and perspectives, fostering an inclusive workplace that represents the global community we engage. We are dedicated to celebrating diverse voices and creating an environment where all are welcomed and valued.
Minimum education
Bachelor's Degree