Measure the job
Build a small representative AI test set
Test a real range of ordinary, difficult, incomplete, ambiguous, and prohibited examples without turning a demo into proof.
What this guide should leave behind
A reusable test set and review rubric expose where the technology helps, fails, refuses, overreaches, or creates additional work.
Work the evaluation in this order
- 01
Define the accepted output and the reviewer. Use synthetic or approved examples that represent the actual range of the job.
- 02
Include normal cases, edge cases, missing information, conflicting instructions, unsafe requests, prohibited data, and examples where the right response is to stop or ask.
- 03
Use the same instructions, configuration, and scoring rubric for each candidate. Record model or product version and test date when visible.
- 04
Review errors by consequence and recovery effort, not only by average score. Retest after material provider, model, prompt, connector, or workflow changes.
Evidence worth seeing
- The test set reflects the job rather than a polished vendor demonstration.
- The reviewer can distinguish acceptable variation from a material error.
- Failures, refusals, unsupported claims, and review time are recorded.
- No result is presented as a general benchmark or independent product certification.
Know when general guidance stops
Use qualified domain and evaluation expertise when a wrong output could cause material financial, legal, health, safety, employment, discrimination, privacy, or customer harm.