- Experience
- 5+ yrs
- Salary
- USD 160,000 – USD 260,000 / year
- Openings
- 1
- Posted
- 1 day ago
- Work mode
- Work from home
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Cohere
Cohere leads as a security-driven enterprise AI company, developing cutting-edge foundational AI models and comprehensive products aimed at resolving realistic business challenges. Our mission centers around training and implementing state-of-the-art AI models to empower enterprises building AI-driven systems. Each team member contributes to advancing our models' capabilities and enhancing customer value, working alongside researchers, engineers, and designers passionate about their craft. Headquartered in Toronto, with offices across global tech hubs including Montreal, London, New York, San Francisco, Paris, Berlin, and Seoul, we invite you to join our innovative team.
Role Overview
We seek a Site Reliability Engineer to join our Model Serving team, dedicated to creating, deploying, and managing the infrastructure supporting Cohere's large language models via accessible API endpoints. This role involves collaborating across teams to deploy optimized NLP models ensuring low latency, high throughput, and strong availability. Opportunities include customer engagement and tailoring deployments to unique requirements.
Key Responsibilities
- Design and implement self-service automation tools for management, deployment, and operation of services.
- Develop custom Kubernetes operators to facilitate language model deployments.
- Automate monitoring and resilience systems, enabling developers to efficiently diagnose and fix issues.
- Ensure adherence to defined Service Level Objectives (SLOs), including participation in on-call rotations.
- Foster strong collaboration with internal developers and help guide Infrastructure team priorities based on their input.
- Contribute to team growth through knowledge sharing and code review practices.
Candidate Profile
- Minimum of five years’ experience managing production infrastructure at scale.
- Expertise designing robust, highly available distributed systems leveraging Kubernetes and managing GPU workloads within clusters.
- Proficiency in developing, coding, and supporting Kubernetes in both development and production environments.
- Experience working with multiple cloud platforms such as GCP, Azure, AWS, OCI, including multi-cloud and hybrid configurations.
- Strong background in Linux-based environments encompassing deployment, support, and troubleshooting.
- Knowledge of managing compute, storage, and network resources, including cost optimization.
- Proven collaboration and problem-solving skills essential to maintain mission-critical systems and facilitate effective teamwork.
- Resilience and adaptability to address evolving technical challenges on a daily basis.
- Understanding of accelerator hardware (GPUs, TPUs, custom accelerators) and their impact on inference latency and throughput.
- Solid comprehension or practical experience with distributed systems architecture.
- Programming experience in Golang, C++, or other high-performance, scalable server languages.
Employee Benefits
- Weekly lunch allowance equivalent to $75 (or £75) in local currency.
- Comprehensive health and dental coverage, supplemented with a dedicated mental health budget.
- Retirement savings plans including RRSP matching, 401K, or pension schemes based on location.
- Parental leave with full salary top-up for up to six months, applicable to any parent.
- Annual enrichment stipends covering arts and culture, fitness, quality time, workspace upgrades, and education for conferences, courses, or coaching.
- Six weeks (30 working days) of paid vacation.
- Travel allowance for visiting other offices and an annual company-wide offsite event.
Work Environment
- A remote-friendly company with offices in major cities worldwide.
- Office amenities include daily lunch programs, snacks, and frequent community activities.
- Remote employees receive co-working space benefits and a $500 stipend for home office setup.
Additional Information
We encourage applications from all backgrounds and prioritize an inclusive work environment. Accommodations for recruitment can be requested to ensure accessibility. AI-driven tools may assist in candidate screening but do not restrict human review. Official communications come only from authenticated Cohere company emails. Beware of fraudulent solicitations requesting payments or services related to hiring.
Compensation
The salary for this full-time position ranges between $160,000 and $260,000 annually.