- Experience
- 3+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 2 weeks ago
- Work mode
- In office
- Education
- Bachelor’s degree or above
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Job Summary
We are looking for an Embedded LLM Systems Engineer responsible for designing, developing, and optimizing large language model (LLM) inference solutions tailored to embedded, mobile, and edge computing devices. The role involves enhancing inference engines, improving model compression, optimizing heterogeneous hardware, and facilitating effective model deployment.
Primary Responsibilities
- Create and improve LLM inference engines for embedded, mobile, and edge platforms.
- Handle operator development, graph optimization, memory management, and integration across multiple backends.
- Utilize frameworks like llama.cpp, TensorRT-LLM, MNN, ONNX Runtime, or similar technologies to develop solutions.
- Research and implement quantization strategies including INT4, INT8, and FP16 formats.
- Leverage relevant computing platforms such as NEON/SVE, Vulkan Compute, or OpenCL for acceleration.
- Perform training and inference consistency checks and aid in streamlined deployment on both cloud and edge environments.
- Explore and apply practical methods to integrate new AI capabilities into embedded products.
Required Qualifications
- Minimum three years of experience in on-device inference, AI infrastructure, embedded systems engineering, or equivalent fields.
- Bachelor’s degree or higher in Computer Science, Electrical/Electronic Engineering, Mathematics, or related disciplines, or equivalent hands-on experience.
- Competency in written and spoken English to understand technical literature, engage in discussions, and produce comprehensive technical documentation.
- Strong command of modern C++ programming focusing on memory models, concurrency, and low-level performance tuning.
- Proficient in Python for tasks such as model conversion, evaluation, scripting, and supporting training tools.
- Experience with CUDA, MediaPipe, or similar technologies is a plus.
Preferred Skills and Experience
- Contributions to open-source inference projects such as llama.cpp, vLLM, TensorRT-LLM, MLC-LLM, or MNN.
- Published research on efficient inference, model compression, or on-device AI deployment in notable conferences or journals.
- Achievements in competitions like ACM-ICPC, NOI, Kaggle, or specialized AI challenges.
- Familiarity with prompt engineering, Retrieval-Augmented Generation, or AI agent frameworks including LangChain and LlamaIndex.
- Hands-on deployment and optimization experience with inference frameworks like vLLM, TGI, llama.cpp, TensorRT-LLM, or MLC-LLM.
- Knowledge about advanced model tuning and alignment techniques such as RLHF, SFT, and DPO.
- Practical expertise in fine-tuning or assessing models using authorized or licensed datasets.
- Strong enthusiasm for emerging LLM technologies and the capacity to apply innovative features into real-world embedded product contexts.
Minimum education
Bachelor's Degree
Skills
Languages
Servicenow