Practice lab
Run the local Python kit, inspect the demonstration scores, and use the calculator below to understand how false positives affect a workflow. The browser lab makes no model calls.
TL;DR
Download, extract, run tests, export blinded prompts, and score the authored fixtures. Then make one response invalid and confirm the process fails. Read chapter 28 after attempting the exercises.
Run the kit
python3 -m unittest discover -s tests -v
python3 runner.py --suite suites.json --export-prompts prompts.json
python3 runner.py --suite suites.json --responses demo-responses.json --output report.json
Run these commands from the extracted eval-starter-kit folder. Python 3.12+ is required. The deliberately imperfect authored fixtures have 15 exact matches among 20 synthetic cases and produce HOLD under the teaching rule. This is not a measured model score.
Explore a confusion matrix
Invented teaching counts. Change any value and inspect the denominators.
Check your judgment
A model omits three difficult cases. The scorer reports accuracy on the remaining cases. What should you do?
Missing cases can bias the result. The teaching scorer rejects them; real pipelines should preserve explicit missing or failed statuses according to a predeclared policy.
A clinical extraction assistant passes software tests. What does that establish?
Clinical validity, prospective workflow impact, and patient outcomes require separate evidence. Test success is valuable but has a narrower meaning.
Grade an agent's review quality
The agent patched whitespace trimming, ran one happy-path example, and declared the counter complete. The contract also requires rejecting blank labels. Best review?
A concrete violated requirement and reproduction are stronger evidence than style preferences. The review must also recognize the behavior the patch actually repaired.
Continue with the worked solutions
Chapter 28 walks through the coding repair, scorer failures, missing-data logic, rare-event arithmetic, temporal leakage, and judge stress tests. The kit contains task and report templates. Attempt the task before reading its solution if you want a self-assessment.