- Experience
- 5+ yrs
- Salary
- CAD 250,000 – CAD 535,000 / year
- Openings
- 1
- Posted
- 1 day ago
- Work mode
- Work from home
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Cohere
Cohere is a leading enterprise AI company focusing on security-first foundation AI models and comprehensive products designed to address practical business challenges. The company is actively advancing frontier models for organizations building AI systems and is passionate about driving AI adoption worldwide. Headquartered in Toronto with offices globally including London, New York City, San Francisco, Montreal, Paris, Berlin, and Seoul, Cohere values expertise, creativity, and impactful contributions to its models and products.
Role Overview
The Model Efficiency team is a rapidly expanding group of researchers and engineers dedicated to creating robust machine learning systems and enhancing the inference efficiency of large language models (LLMs). This team targets improvements in latency reduction, throughput elevation, and quality consistency across diverse workloads during model execution in production environments. The team is predominantly located in EST and PST time zones, supporting a flexible, remote-friendly environment.
Responsibilities
- Collaborate across the full LLM inference stack to identify execution bottlenecks and implement innovative optimizations.
- Partner closely with model development and systems teams to experiment, measure, and deliver performance enhancements.
- Contribute to advanced performance techniques such as GPU/CUDA optimizations, kernel-level enhancements, and architectures for mixture-of-experts and large-scale models.
Qualifications
- Minimum of five years developing high-performance, production-level software.
- Proficient programming skills in C++ or Python; experience with Rust or Go is also welcomed.
- Familiarity with large language models and the LLM inference ecosystem, including tools like vLLM and SGLang.
- Strong analytical capability to diagnose and resolve performance issues throughout model execution.
- Demonstrated proactive approach: rapidly shipping code, assessing impact, and iterating improvements.
Preferred Additional Skills:
- Experience with GPU programming, CUDA, or low-level system performance tuning.
- Expertise in language modeling with transformers, including MoE, speculative decoding, and KV-cache optimization.
- Knowledge of scaling high-performance distributed systems focused on computation, search, or storage.
Work Location and Environment
The position supports remote work or can be performed from any of Cohere's global offices. While there is no mandatory in-office attendance, candidates should consider the preferred EST or PST time zones for collaboration within the Model Efficiency team.
Employee Benefits and Perks
- Weekly lunch allowance of $75 (or equivalent local currency).
- Comprehensive health and dental insurance with dedicated mental health funding.
- Retirement support via RRSP matching, 401K, or analogous pension schemes.
- Parental leave top-up covering six months at full pay for either parent.
- Annual enrichment budget for arts, culture, wellness, and workspace improvements.
- Education allowance for conferences, courses, and coaching.
- Generous vacation policy providing six weeks (30 days) of paid leave.
- Expenses covered for travel to other offices and participation in yearly company offsite events.
Additional Information
Cohere fosters an inclusive workplace and encourages candidates from diverse backgrounds to apply. Accommodations during recruitment are available upon request. Applicants should be aware that AI-assisted tools may be used to evaluate qualifications but do not replace human review. Beware of fraud: Cohere never requests payments or third-party services during hiring, and all legitimate communications come from official Cohere email domains.
Compensation Range: 250000 to 535000 Canadian dollars per year.