AI Research Engineer - Frontier Models & Benchmarking
Dublin, County Dublin, Ireland · Full Time
Be the first to apply
- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 3 days ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Role
Join a pioneering AI company in Dublin focused on creating training and evaluation platforms for world-leading AI labs. This role is crafted for specialists deeply involved in frontier AI models, particularly Large Language Models (LLMs). Unlike common roles focusing on application-level tasks, this position demands expertise in probing, testing, and benchmarking the models themselves.
Key Responsibilities
- Design and develop benchmarking tools tailored to cutting-edge LLMs.
- Analyze and compare the performance characteristics and behaviors across various AI models.
- Create challenging tasks and datasets that reveal strengths and vulnerabilities of models.
- Conduct adversarial testing to identify breaking points and failure modes.
- Establish and apply rigorous evaluation methodologies and metrics.
- Work with synthetically generated training and evaluation data.
- Investigate model reasoning, reliability, and unexpected responses.
- Engage in post-training procedures including reinforcement learning experiments.
- Orchestrate experiments encompassing multiple frontier AI models.
Desired Experience and Expertise
- Proven experience constructing or contributing to recognized/public LLM benchmarking initiatives.
- Expertise in benchmarking LLM capabilities or performing adversarial testing and red teaming.
- Background in AI or model safety research.
- Experience with post-training, reinforcement learning, or synthetic data generation for model training and evaluation.
- Design and implementation of evaluation frameworks, metrics, or automated judges for models.
- Research into the behavior, failures, and nuances of LLMs.
- Familiarity working across diverse models like Claude, GPT, Gemini, Llama, or DeepSeek.
Qualifications and Attributes
Strong research background is highly preferred, including PhD or substantial research output with publications in prestigious conferences (e.g., NeurIPS, ICML, ICLR) or demonstrable achievements in building benchmarks and evaluation platforms.
The ideal candidate is someone who has actively conducted deep model exploration rather than merely using buzzwords or standard tools.
Exclusions
If your expertise is mostly in:
- Retriever-Augmented Generation (RAG)
- LangChain frameworks
- Vector databases
- Prompt engineering
- Chatbot development or agent orchestration
- General use of LLM APIs or adding GenAI features to existing applications
this role might not be suitable as it demands deeper benchmarking and model evaluation experience.
Working Conditions
This is a full-time position requiring presence five days a week onsite in central Dublin. The team is small, energetic, and fast-paced, ideal for individuals suited to a startup-like environment rather than a conventional 9-to-5 job.
Additional Information
If you regularly engage with new frontier models by testing, breaking, and comparing their behaviors, this opportunity is tailored for you.
Minimum education
Doctorate