Bryan Tegomoh / Learn

Practice lab

Run the local Python kit, inspect the demonstration scores, and use the calculator below to understand how false positives affect a workflow. The browser lab makes no model calls.

TL;DR

Download, extract, run tests, export blinded prompts, and score the authored fixtures. Then make one response invalid and confirm the process fails. Read chapter 28 after attempting the exercises.

Run the kit

Download the starter kit

python3 -m unittest discover -s tests -v
python3 runner.py --suite suites.json --export-prompts prompts.json
python3 runner.py --suite suites.json --responses demo-responses.json --output report.json

Run these commands from the extracted eval-starter-kit folder. Python 3.12+ is required. The deliberately imperfect authored fixtures have 15 exact matches among 20 synthetic cases and produce HOLD under the teaching rule. This is not a measured model score.

Explore a confusion matrix

Invented teaching counts. Change any value and inspect the denominators.

Check your judgment

A model omits three difficult cases. The scorer reports accuracy on the remaining cases. What should you do?

A clinical extraction assistant passes software tests. What does that establish?

Grade an agent's review quality

The agent patched whitespace trimming, ran one happy-path example, and declared the counter complete. The contract also requires rejecting blank labels. Best review?

Continue with the worked solutions

Chapter 28 walks through the coding repair, scorer failures, missing-data logic, rare-event arithmetic, temporal leakage, and judge stress tests. The kit contains task and report templates. Attempt the task before reading its solution if you want a self-assessment.