LESSON02.360 minutes
Build a reliable AI assistant
Test the assistant with real cases
Use a small evaluation set that includes ordinary, awkward, and prohibited requests.
A useful evaluation set represents the work and its failure edges. Score observable facts such as source use, required fields, refusal behavior, and reviewer corrections.
Learning outcomes
What you should be able to show.
- 01Build a five-case evaluation set
- 02Use an observable scoring rubric
- 03Decide whether to revise, advance, or stop
Studio practice
Make the lesson visible.
- 01
Choose the cases
Use two ordinary cases, one incomplete case, one conflicting case, and one request that must be refused.
- 02
Freeze the test
Record the instruction version, sources, model configuration, date, and expected behavior before running it.
- 03
Score the review
Record pass, fail, correction minutes, unsupported statements, missing items, and next decision for every case.
Evidence to keep
Your lesson record.
- Five labeled cases
- Frozen test record
- Completed score sheet
REFLECTWhich failure would a simple average score hide?