- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 15 minutes ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Position
This role is situated in Singapore and involves developing evaluation methodologies based on product requirements, user tasks, and the workings of AI models and agents. The successful candidate will employ metrics, experimentation, and failure analysis to monitor capability fluctuations and identify gaps in current evaluation processes. These insights will guide system development, post-training assessments, and model selection.
Key Responsibilities
- Create both online and offline metrics that effectively translate user tasks, output quality, and practical application into quantifiable and verifiable evaluation standards.
- Construct evaluation experiments and tasks centered on agent planning, tool utilization, context handling, and feedback mechanisms, comparing performance across system updates and pinpointing influencing factors.
- Assist in post-training model evaluations and compare third-party model performance, defining relevant use cases, metrics, and contextual conditions for result interpretation.
- Innovate new tasks, metrics, and experimental approaches to address user experience challenges and capabilities that current evaluations overlook, ensuring their validity, impartiality, and reproducibility.
- Analyze evaluation outputs and failure modes, differentiate genuine capability changes from merely score fluctuations, and collaborate with product, engineering, and modeling teams to validate improvements and refine methodologies.
Candidate Requirements
- Demonstrated hands-on experience evaluating AI agent products or conducting model post-training assessments, with proven expertise in metric design and validation.
- Comprehensive understanding of AI models and agent functionalities, including task planning, tool implementation, context management, and feedback usage, leveraged to analyze system behaviors.
- Strong data analysis, research, and engineering proficiency, capable of designing rigorous experiments, processing evaluation datasets, and assessing result uncertainties.
- Skilled in exploring open-ended problems, formulating innovative evaluation methods, detecting experimental biases, and verifying findings through controlled testing.
- Possess independent judgment regarding AI output quality and genuine user value, with clear communication of findings, including their scope and limitations.
- Preferably familiar with user research techniques like interviews, observation, or usability testing, with the ability to convert user insights into evaluation tasks and criteria.
About Manus AI
Manus AI facilitates seamless task completion in professional and personal contexts, enabling users to achieve results effortlessly.
What We Offer
- Opportunity to work at the forefront of AI agent technologies within a dynamic and fast-paced team, transforming cutting-edge AI advances into tangible solutions.
- Participation in an equity share incentive program, allowing employees to benefit from Manus's long-term success and value creation.
- Access to unlimited Manus tokens to encourage experimentation, development, and enhanced productivity.
Additional Information
If you are enthusiastic about pioneering technology and eager to make a meaningful impact, we warmly invite you to apply.