- Experience
- 6+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 11 hours ago
- Work mode
- In office
- Education
- Bachelor’s or Master’s degree in Computer Science or related field
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Role Overview
We are seeking a seasoned AI Evaluation Engineer to take charge of designing, calibrating, and managing evaluation frameworks and gate thresholds specific to various agent archetypes. This pivotal senior technical role involves leading gate reviews under team leadership, managing regression testing methods for AI/ML model modifications, and ensuring the maintenance and quality oversight of evaluation processes.
Key Responsibilities
- Create and sustain offline evaluation tools including golden datasets, regression test packs, and adversarial/safety probes alongside managing scoring pipelines for different agent categories.
- Lead operability assessments by reviewing evidence packages, reproducing evaluation outcomes, and making informed go/no-go decisions with detailed documentation.
- Oversee the regression test methodology for model updates and collaborate with AI/ML operations to adjudicate results against established benchmarks.
- Adjust gate thresholds to align with live production data and uphold evaluation integrity by rotating datasets and calibrating judges.
- Compile monthly reports detailing quality metrics, evaluation trends, categorization of failure modes, and documented instances of defects with replication steps.
- Provide mentorship to team members and ensure high standards of quality and consistency in their deliverables.
Qualifications and Skills
- University degree (Bachelor’s or Master’s) in Computer Science or equivalent discipline.
- At least six years’ experience in machine learning, data science, or software development with a strong emphasis on evaluation or quality assurance.
- Proven hands-on expertise in assessing large language models (LLMs) or machine learning systems.
- Competence in designing LLM/agent evaluation strategies with statistical depth.
- Proficiency with Python programming and evaluation frameworks such as promptfoo, DeepEval, or other custom tools.
- Strong data analysis capabilities and metric interpretation skills.
- Familiarity with tracing and observability tools.
- Meticulous attention to detail and analytical mindset.
- Excellent technical communication skills for preparing gate review documentation.
- Capability to influence development teams on quality matters.
- Understanding of telecommunications customer intentions and journey mapping.
Skills
How they work
Communication
Problem Solving
Attention to Detail
Leadership
Negotiation