Zohreh Razavi
← back to work

AI product team · 2026

LLM Evaluation & Quality Harness

An evaluation harness that scores LLM answers for accuracy, safety, and groundedness before each release.

Client
AI product team
Role
Project Manager
Timeline
Apr 2026 – Aug 2026
Domain
AI / Quality
Approach
Agile
Team size
5–9
Budget
Up to $250K

Objective

Deliver a repeatable evaluation harness so teams can score model and prompt changes against a shared quality bar before shipping.

What I did

  • ▸ Defined quality criteria and release gates with product and AI leads
  • ▸ Coordinated dataset design, scoring jobs, and reporting
  • ▸ Planned sprints and tracked eval coverage across features
  • ▸ Set up regression alerts when quality dropped
  • ▸ Rolled the harness into the existing release process

Key deliverables

  • ▸ Evaluation datasets
  • ▸ Scoring pipeline
  • ▸ Quality dashboards
  • ▸ Release gates
  • ▸ Team rollout

Outcome

Gave the team a shared quality bar, catching regressions before they reached users.

Tools & tech

  • LLMs
  • Evals
  • Python
  • CI
← back to work