- Experience
- 5+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 1 week ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Role Overview
As an Inference Performance Engineer, you will be responsible for the efficiency and performance of our inference infrastructure. Your efforts will directly influence how we serve machine learning models efficiently under varying workloads, changing traffic, and diverse hardware setups. Collaborating closely with fleet operations engineers, you will manage critical performance aspects such as caching, batching, quantization, decoding, and kernel-level tuning to enhance throughput and reduce latency while maintaining system reliability and model accuracy.
Key Responsibilities
- Enhance throughput, reduce costs, and minimize tail latency using strategies like KV-cache management, ongoing batching, speculative decoding, and quantization techniques.
- Optimize prefill and decode operations for long context workloads backed by actual production traffic data.
- Fine-tune request routing across internal infrastructure and external providers by evaluating cost, capacity, and performance metrics.
- Engage deeply with serving engines such as vLLM, SGLang, and TensorRT-LLM, including working below the framework layer when necessary.
- Develop tools for profiling and measurement to detect computational, memory, and timing bottlenecks.
Candidate Profile
- 5 or more years of experience in machine learning systems, inference infrastructure, or performance engineering demonstrating tangible gains in cost or latency improvements.
- In-depth expertise in model serving concepts including prefill and decode stages, memory bandwidth optimization, batching, and handling concurrency.
- Practical knowledge of serving platforms like vLLM, SGLang, or TensorRT-LLM in production environments.
- Proficiency in Python, along with competence in systems programming languages such as C++ or Rust.
- Strong understanding of GPU performance factors including CUDA programming, NCCL, mixed-precision arithmetic, kernel optimization, memory layout, and quantization.
Work Culture and Values
We value team players who make collaborative work enjoyable and embrace innovative, daring ideas. Adaptability is crucial; we welcome applicants who may not meet every criterion but show a willingness to learn and grow.
About the Company
Our company focuses on developing intelligent AI systems that continuously evolve rather than remain static. We strive to create flexible, personalized, and accessible AI powered by efficient computing, which broadens access and democratizes innovation. We foster a high concentration of highly motivated individuals passionate about pioneering continual adaptation in intelligence technology.
Employee Benefits
- Flexible working arrangements with onsite collaboration in the Bay Area, a distributed global team, and regular team offsites.
- Annual travel stipend called the Adaption Passport, encouraging exploration of new countries to foster personal and professional growth.
- Weekly meal allowance to support take-out or grocery delivery options.
- Comprehensive health insurance and generous paid vacation policies promoting well-being.