- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 3 weeks ago
- Work mode
- In office
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Role
We are developing an evaluation benchmark for advanced AI browsing agents. Your responsibility is to craft research challenges that a cutting-edge AI cannot solve, even when given complete web access and multiple attempts. This position focuses on investigative research rather than subject-matter expertise or content writing. You will begin with a verifiable fact and reverse-engineer it into a question that makes it extremely difficult to find, supported by a thorough and auditable evidence trail.
Deliverables
- Create natural-language research questions with concise, stable, and objectively verifiable answers.
- Provide independent clues that can be checked separately, covering various fact types such as dates, people, places, organizations, works, events, records, and quantities, each with specific requirements.
- Produce a validation record documenting the obvious searches performed and their outcomes.
Required Qualifications
- Proven skill in open-web research, including locating primary documents and accessing government and institutional databases, archives, registries, and PDF files.
- Precision in sourcing, citing exact pages, tables, and sections rather than general homepages.
- Ability to research unfamiliar topics independently.
- Native or near-native proficiency in written English.
- Strong aptitude for creating detailed structured documentation; the evidence trail constitutes most of the task.
- Experience with large language model (LLM) evaluation, red-teaming, or building benchmarks.
- Background in one or more of the following areas: reference librarianship, archival research, special collections, investigative journalism, professional fact-checking, OSINT, due diligence, KYC, investigative research, patent or prior-art searching, legal discovery searches, genealogy, or developing competitive quizzes and puzzle hunts.
Preferred Qualifications
- Additional experience with LLM evaluation, red-teaming, or benchmark creation.
- Knowledge of JSON and structured data formats.
How they work
Problem Solving
Attention to Detail
Motivation
Languages
English