Senior Machine Learning Engineer (Large Systems)
Cambridge, England, United Kingdom · Full Time
Be the first to apply
- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 6 days ago
- Work mode
- In office
- Education
- Bachelor's or higher in Machine Learning, Computer Science, Mathematics, Data Science, or a related discipline
- Eligibility
- Applicants must have the legal right to work in the United Kingdom. Visa sponsorship or assistance is not available.
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Graphcore
Graphcore specializes in pioneering the next generation of AI computing technology. Our team consists of experts in semiconductors, software development, and artificial intelligence who collaborate to build an end-to-end AI computing platform—ranging from hardware to data center-scale infrastructure. Backed by SoftBank Group’s substantial long-term investments, our goal is to deliver cutting-edge technologies into SoftBank’s rapidly growing AI ecosystem. We are expanding globally to assemble talented individuals passionate about shaping the future of AI.
Role Overview
The Senior Machine Learning Engineer will be part of the Applied AI team, focusing on advancing AI technologies by developing and optimizing machine learning models designed specifically for Graphcore’s unique hardware. The role requires working at scale on performance-critical systems. Collaborating closely with both software development and research teams, you will identify innovation opportunities to differentiate our technology in the marketplace. We seek highly skilled engineers experienced in large-scale AI model implementation motivated by creating impactful technological improvements.
Team Mission
The Applied AI team represents our customers' interests by mastering current AI models, software, and applications to ensure flawless integration and scalability of Graphcore's technology within the AI ecosystem. We create reference applications, optimize key software libraries (such as kernels tailored for our hardware), and work hand-in-hand with Research to publish novel work on topics including efficient computation, model scaling, and distributed training and inference across various AI modalities.
Key Responsibilities
- Implement state-of-the-art machine learning models, optimizing them for speed and accuracy, scaling up to thousands of accelerators.
- Test and assess new internal software releases, submit constructive feedback to engineering teams, perform code fixes, and participate in code reviews.
- Conduct benchmarking of models and key machine learning methods to detect bottlenecks and increase model efficiency.
- Design and carry out experiments involving innovative AI techniques, analyze outcomes, and refine methods.
- Collaborate extensively with Research, Software, and Product teams to help develop and validate the next generation of Graphcore AI hardware.
- Maintain engagement with the AI research community and stay informed on the latest advancements in the field.
Candidate Profile
- Essential qualifications include a Bachelor’s, Master’s, PhD, or equivalent experience in Machine Learning, Computer Science, Mathematics, Data Science, or related disciplines.
- Proficiency in deep learning frameworks such as PyTorch and JAX.
- Strong programming skills in Python and C++.
- Comprehensive experience with deep learning spanning model development, training, optimization, and evaluation.
- Solid capability in designing, executing, and reporting machine learning experiments.
- Well-developed understanding of identifying and mitigating performance bottlenecks.
- Ability to adapt swiftly in a fast-moving environment.
- Enjoy collaborative, cross-functional teamwork.
- Excellent communication skills, able to articulate complex technical material to varied audiences.
- Desirable experiences include MLOps for Kubernetes clusters, building production systems involving large language models, expertise in efficient computing employing low-precision arithmetic, and proficiency in writing C++, Triton, or CUDA kernels to optimize ML model performance.
- Experience with distributed training or inference on clusters of 64 or more accelerators.
- Familiarity with HPC systems and network technologies such as Infiniband, NVLink, RoCE.
- A history of contributions to open-source projects or research publications in relevant AI fields.
- Knowledge of cloud computing platforms.
- Motivation to actively present, publish, and participate in AI community events.
Compensation and Benefits
Graphcore offers a competitive salary package including flexible work arrangements, generous annual leave, private medical and dental insurance, a pension scheme with up to 5% company match, life assurance, and income protection. Our parental leave policy supports family needs, and an employee assistance program provides health, mental well-being, and bereavement support. The central Bristol office features healthy snacks and a barista bar. We are committed to diversity and inclusivity in our workforce and provide reasonable accommodations during the hiring process.
Additional Information
Applicants must have the legal right to work in the United Kingdom; unfortunately, we cannot offer visa sponsorship or support at this time.
Level
Senior
Minimum education
Bachelor's Degree