Senior Research Scientist | Model Steering
Cologne, North Rhine-Westphalia, Germany · Full Time
Be the first to apply
- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 1 week ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About DeepL
DeepL is an innovative AI product and research firm dedicated to delivering secure and intelligent solutions to complex challenges in business communication. Trusted by over 200,000 companies and millions of users worldwide, DeepL's Language AI platform offers superior translation, enhanced writing, and real-time voice translation across 228 markets globally. Established in 2017 by Jaroslaw "Jarek" Kutylowski, DeepL has grown to nearly 1,000 employees and enjoys support from prestigious investors such as Benchmark, IVP, and Index Ventures.
DeepL aims to become the foremost authority in trusted AI technology by developing products that improve communication, nurture connections, and make a significant impact. This vision is supported by a passionate team of innovators and researchers committed to advancing AI for simplifying and enriching human collaboration.
Team and Mission
The Language AI teams are core to DeepL’s achievements, working collaboratively with engineers, product managers, and designers to create world-leading language AI systems for flawless translation in demanding environments. Team members take full ownership of the machine learning model lifecycle, including data handling, training processes, quality assurance, and operational integration, enabling broad influence across the company.
Key Responsibilities
- Lead the development of steerable translation models tailored to user preferences, rules, and contextual inputs.
- Conduct hands-on research and apply post-training techniques such as supervised fine-tuning, knowledge distillation, and reinforcement learning methods that optimize translation quality.
- Design and implement reward and evaluation models for translation assessment, apply grading rubrics, and address challenges such as reward gaming and estimation errors.
- Pursue the advancement of models capable of processing multimodal content for enhanced translation quality.
- Manage the entire model development cycle including prototyping, experimentation, training, evaluation, fine-tuning, and deploying into real-time production systems.
- Establish rigorous protocols for model evaluation, reproducibility, monitoring, and continuous improvement in operational environments.
- Guide, mentor, and collaborate with researchers and engineers to elevate overall model performance standards.
Candidate Profile and Qualifications
- Proven expertise in making large language models steerable using methods such as instruction tuning, latent space manipulations, steering vectors, or constrained decoding/encoding.
- Deep practical experience in large language model post-training strategies including supervised fine-tuning, direct preference optimization, knowledge distillation, or reinforcement learning techniques like RLHF/RLAIF, PPO/GSPO.
- Strong data-centric skills to develop synthetic and preference-data pipelines, utilizing LLM-based judgment, efficient data curation, and analysis of data mixtures.
- Proficiency in creating evaluation metrics and reward functions incorporating automatic measures, human-in-the-loop assessments, and robust validation strategies.
- Hands-on experience with model training, large-scale experiments, debugging complex ML pipelines, and integrating machine learning solutions in production, with a focus on real-world application impact.
- Track record of ownership and delivering substantial research projects with strong execution skills and prior mentorship experience.
- Advanced programming and experimentation proficiency in Python, leveraging frameworks such as PyTorch, JAX, or TensorFlow, combined with effective communication aligning research priorities with product teams.
Preferred Additional Experience
- Experience training large-scale models with distributed and multi-node training frameworks such as FSDP, DeepSpeed, or Megatron-like systems.
- Skills in fine-tuning large reasoning models for specific behaviors without compromising their inference abilities.
- Background in machine translation, multilingual natural language processing, or language quality estimation research.
- Familiarity with scalable inference and serving technologies including vLLM, SGLang, TensorRT-LLM, and expertise in long-context model handling.
- Contributions to top-tier research publications.
Benefits and Work Culture
- Join a diverse international workforce representing over 90 nationalities, with an expanding global presence across multiple countries.
- Open communication culture with regular, constructive feedback to support growth and collaboration built on empathy and mutual respect.
- Hybrid working model allowing office presence twice weekly combined with flexible work hours to accommodate team time zones and promote balance.
- Equity participation through Virtual Shares to reward individual contributions aligned with company growth.
- Frequent in-person events fostering team bonding, from local gatherings to comprehensive company-wide meetups.
- Monthly "Hack Friday" sessions encouraging innovation and cross-team project collaboration.
- Generous annual leave of 30 days excluding public holidays, complemented by mental health resources for well-being support.
- Competitive benefits package tailored geographically to meet diverse employee needs worldwide.
Equal Opportunity and Inclusion
DeepL is committed to equal opportunity employment and embraces diversity, encouraging applicants from all backgrounds to bring their authentic selves to work. We believe diverse perspectives strengthen our innovation and shared mission to break down global language barriers.
Level
Senior
Industry
Artificial Intelligence