- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 1 day ago
- Work mode
- In office
- Education
- Master's or PhD
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Company
We are collaborating with a leading global organization specializing in information and communications technology infrastructure and smart devices. They provide comprehensive, full-stack solutions across various scenarios serving carriers, enterprises, governments, and individual consumers worldwide.
Job Overview
This position focuses on innovating high-efficiency compression and inference strategies spanning traditional video/media codecs and contemporary Large Language Model (LLM) inference systems. The role requires designing intelligent processing pipelines that offer enhanced visual quality at reduced bitrates, alongside creating algorithms to minimize memory usage and overcome computational constraints in generative AI deployment.
Key Responsibilities
- Accelerate LLM inference by researching and implementing sophisticated compression approaches emphasizing KV cache optimization, quantization of models, and alleviation of memory bandwidth bottlenecks during autoregressive decoding.
- Develop advanced components of classical video codecs, aiming to boost Rate–Distortion performance, refine entropy coding, and advance quantization methods suitable for practical applications.
- Create and improve AI-powered video coding modules, including AI-based loop filtering, optical flow computation, and smart rate control optimization.
- Facilitate the transition from AI research to real-world deployment by fine-tuning deep learning models for efficient inference and ensuring smooth incorporation of compression algorithms into platforms like vLLM.
- Perform thorough objective and subjective assessments of system performance and quality, employing metrics such as PSNR, VMAF for video, and perplexity, zero-shot tests, latency, and throughput for LLM systems.
Required Qualifications
- Master’s degree or PhD in Computer Science, Electronic Engineering, Mathematics, or closely related fields (PhD preferred).
- Comprehensive knowledge of video coding basics, such as prediction, transform and entropy coding, quantization, backed by practical experience with standards like H.265/HEVC, AV1, or H.266/VVC.
- In-depth understanding of Transformer neural network architectures and attention mechanisms, plus familiarity with key generative AI inference challenges including memory bandwidth limitations.
- Proficiency in Python and C/C++ programming; experience with building, training, and adjusting models using frameworks like PyTorch and TensorFlow.
Preferred Qualifications
- Experience with Image Signal Processing workflows, including demosaicing, denoising, and tone mapping techniques.
- Expertise in image processing enhancements through computer vision, such as de-blurring, artifact cleaning, and high dynamic range (HDR) imaging.
- Familiarity with hardware acceleration methods like SIMD or CUDA to optimize video and tensor operations.
Additional Information
Only shortlisted candidates will be contacted regarding the recruitment process. By submitting your application and CV, you consent to PERSOL Singapore Pte Ltd and its affiliates collecting and using your personal data as per their Privacy Policy.
Minimum education
Master's Degree