Reusable decision kit
Representative test-set sheet
A reusable set of normal, difficult, ambiguous, unsafe, and prohibited examples.
What the completed kit should do
Candidates face the same job-specific evidence instead of different polished demonstrations.
Capture these facts in your approved system
- Test ID, scenario family, approved input, and expected behavior
- Acceptance criteria, reviewer, and consequence of error
- Product, plan, model or version when visible, configuration, and date
- Observed output, refusal, unsupported claim, review effort, and recovery
- Result, severity, follow-up, retest trigger, and decision impact
Use the structure in this order
- 01
Define accepted behavior and create approved examples.
- 02
Run every candidate under the same conditions.
- 03
Score with the same reviewer rubric.
- 04
Investigate consequential failures and preserve the dated result.
Close the evidence loop
The team can explain where each candidate works, fails, stops, and creates review work.
Keep sensitive material out of this site
- A small internal test is not a public benchmark or certification.
- Do not use real sensitive records without explicit qualified approval.
- Retest after material changes to models, instructions, connectors, data, or workflow.