Senior Research Scientist - Model Steering
Munich, Bavaria, Germany · Full Time
Be the first to apply
- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 1 week ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Overview
DeepL, established in 2017 and employing approximately 1,000 people globally, is a leading AI company specializing in Language AI products. Serving over 200,000 businesses and millions of users in 228 markets, it aims to lead the AI industry by creating trusted and intelligent solutions that simplify communication and foster global connections.
Team and Mission
The Language AI teams at DeepL form the core of the company's success, focusing on designing and managing the full lifecycle of advanced machine learning models tailored for language translation. These teams work collaboratively across research, engineering, product, and design to deliver highly accurate translation systems.
Role Summary
We seek a Senior Research Scientist to oversee fine-tuning, post-training procedures, model steerability, and reinforcement learning for the next generation of DeepL's large language models (LLMs) applied in translation. This influential role involves rapid prototyping, conducting extensive experiments, and spearheading innovations from research through to production deployment.
Key Responsibilities
- Lead creation of translation models adaptable to user preferences, rules, and contextual inputs.
- Conduct hands-on research on post-training approaches including supervised fine-tuning, knowledge distillation, preference optimization, and reinforcement learning focused on enhancing translation quality.
- Develop reward and evaluation models for translations, addressing challenges such as reward exploitation and estimation inaccuracies.
- Advance research into models that incorporate multimodal content and contextual data to boost translation accuracy.
- Manage the end-to-end process of model development, including prototyping, ablation studies, training, evaluation, optimization, and real-time system integration.
- Implement robust practices for evaluation, reproducibility, monitoring, and ongoing improvements in production settings.
- Mentor team members, facilitate collaborative research efforts, and raise standards for translation model quality.
Candidate Profile
- Proven expertise in enabling large models to follow instructions and exhibit steerable behaviors, utilizing methods such as instruction tuning, latent space techniques, and constrained encoding/decoding.
- Substantial experience with LLM post-training methodologies including supervised fine-tuning, direct preference optimization, knowledge distillation, and reinforcement learning algorithms like RLHF, RLAIF, PPO, or GSPO.
- Strong capability in managing data pipelines for synthetic and preference-based datasets, including curation, filtering, and designing data mixtures.
- Experience designing evaluation metrics and reward signals, including automatic methods, human-in-the-loop approaches, and handling non-verifiable reward signals.
- Practical skills in training models, debugging pipelines, conducting experiments, and deploying ML systems with clear product impact.
- Track record of independent research ownership combined with effective mentorship in fast-paced applied research environments.
- Proficiency in programming and experimentation using Python and ML frameworks like PyTorch, JAX, or TensorFlow, accompanied by clear communication and alignment with engineering and product teams.
Desirable Qualifications
- Experience with large-scale fine-tuning and training on distributed systems using frameworks such as FSDP, DeepSpeed, or Megatron.
- Expertise in fine-tuning reasoning models without compromising their reasoning abilities.
- Background in machine translation, multilingual natural language processing, or language quality estimation.
- Familiarity with scalable inference and serving technologies like vLLM, SGLang, TensorRT-LLM, and long-context language modeling.
- Publications in leading academic venues.
Work Environment and Benefits
- Diverse international team comprising over 90 nationalities and a global footprint including UK, Germany, Netherlands, Poland, USA, and Japan.
- An open culture emphasizing transparent communication, constructive feedback, empathy, and a growth mindset.
- Hybrid work model with flexible scheduling; team members come to the office twice per week to encourage connection and spontaneity alongside work-from-home flexibility.
- Equity participation through Virtual Shares to link employee efforts to company success.
- Regular in-person social and team-building activities including business unit gatherings, onboarding events, and company-wide meetings.
- Dedicated monthly innovation days (“Hack Fridays”) fostering creativity and cross-team collaboration.
- Generous paid time off with 30 days of annual leave excluding public holidays and mental health resources.
- Competitive compensation and benefits designed to accommodate the diverse team across various locations.
Equal Opportunity
DeepL values authenticity and inclusiveness, encouraging candidates of all backgrounds to apply and contribute diverse perspectives. The company is committed to providing an environment where everyone can thrive and help eliminate language barriers worldwide.
Level
Senior
Industry
Artificial Intelligence