Singtel

Senior AI Evaluation Engineer

Singtel

Singapore · Full Time

Be the first to apply

Experience
6+ yrs
Salary
Openings
1
Posted
2 days ago
Work mode
In office
Education
Bachelor's or Master's in Computer Science or related field
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About the Role and Organization

Singtel has launched AIDA, a new business unit focused on Artificial Intelligence and Data Analytics, aimed at driving the company's AI transformation with strategic precision. This initiative integrates intelligence into all aspects of the business to enhance human capabilities and unlock new potential. The transformation aligns people, platforms, and processes to foster an AI-literate culture that empowers employees and redefines telecommunications innovation.

Position Summary

As the Senior AI Evaluation Engineer, you will function as the lead technical evaluator responsible for designing, calibrating, and adjudicating evaluation suites and gate thresholds tailored to various agent archetypes. You will conduct gate reviews, manage regression-pack methodologies for AI/ML model changes, and oversee evaluation quality and drift management.

Key Responsibilities

  • Develop and sustain offline evaluation suites such as golden sets, regression packs, adversarial and safety probes, along with continuous evaluation scoring pipelines across different archetypes.
  • Perform operability gate reviews by analyzing evidence packs, reproducing evaluation outcomes, and recommending go/no-go decisions with detailed documentation.
  • Lead the model-update regression pack methodology and adjudicate regression runs against archived baselines collaboratively with the AIML Operations team.
  • Calibrate gate thresholds based on production performance and maintain evaluation drift hygiene, including rotation of golden sets, using hold-out sets, and judge calibration.
  • Generate monthly quality reports per agent covering evaluation trends, failure modes taxonomy, and defect clusters with reproducible traces.
  • Provide mentorship to team members and review their output to ensure consistency and quality.

Required Qualifications and Skills

  • A Bachelor's or Master's degree in Computer Science or a related discipline.
  • At least 6 years of experience in machine learning, data science, or software engineering with a solid focus on evaluation or quality.
  • Practical experience in evaluating Large Language Models (LLM) or machine learning systems.
  • Proficiency in designing LLM/agent evaluation frameworks with strong statistical rigor.
  • Expertise in Python programming and evaluation tooling, such as promptfoo, DeepEval, or customized evaluation harnesses.
  • Skillful in data analysis and interpreting evaluation metrics.
  • Experience with tracing and observability tools.
  • Knowledge of adversarial testing and red teaming techniques.

Preferred Qualifications

  • Experience working with agentic systems or retrieval-augmented generation (RAG) pipelines.
  • Familiarity with telecommunications sector AI applications.
  • Exposure to management-level presentations.
  • Understanding of Responsible AI and safety evaluation methodologies.
  • Insight into telecommunications customer intents and journey mapping.

Soft Skills

  • Strong analytical skills with great attention to detail.
  • Clear and effective technical writing for reporting gate evaluations.
  • Ability to influence engineering teams on quality standards and improvements.

Join Us

If you are ready to explore substantial opportunities and impact future innovations, join Singtel and accelerate your career through meaningful projects, ongoing learning, and measurable results.

Minimum education

Master's Degree

How they work

Communication Problem Solving Attention to Detail Leadership Persuasion & Influence

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer