Skip to content

Measure the job

Build a small representative AI test set

Test a real range of ordinary, difficult, incomplete, ambiguous, and prohibited examples without turning a demo into proof.

TARGET OUTPUT

What this guide should leave behind

A reusable test set and review rubric expose where the technology helps, fails, refuses, overreaches, or creates additional work.

SEQUENCE / 04

Work the evaluation in this order

  1. 01

    Define the accepted output and the reviewer. Use synthetic or approved examples that represent the actual range of the job.

  2. 02

    Include normal cases, edge cases, missing information, conflicting instructions, unsafe requests, prohibited data, and examples where the right response is to stop or ask.

  3. 03

    Use the same instructions, configuration, and scoring rubric for each candidate. Record model or product version and test date when visible.

  4. 04

    Review errors by consequence and recovery effort, not only by average score. Retest after material provider, model, prompt, connector, or workflow changes.

CHECK BEFORE GATE

Evidence worth seeing

  • The test set reflects the job rather than a polished vendor demonstration.
  • The reviewer can distinguish acceptable variation from a material error.
  • Failures, refusals, unsupported claims, and review time are recorded.
  • No result is presented as a general benchmark or independent product certification.
ESCALATION BOUNDARY

Know when general guidance stops

Use qualified domain and evaluation expertise when a wrong output could cause material financial, legal, health, safety, employment, discrimination, privacy, or customer harm.