Bryan Tegomoh / Learn
Chapter 05

Design an evaluation before building it

TL;DR

Define the decision, task, population, inputs, outputs, reference, and failure criteria before building. Match the strength of your conclusion to the construct actually tested.

An eval should test an explicit claim. “This model is good at medicine” is too broad. “Under the supplied synthetic reconciliation policy, this configured system preserves medication status and does not invent missing dose fields” is testable.

The evaluation contract

Use the downloadable task-contract template. Complete every field before implementation:

Field Example
Decision Whether to advance an extraction assistant to a supervised pilot
Intended user A clinician reviewing a draft, not an autonomous prescribing system
Unit One encounter note and its structured output
Construct Faithful extraction of stated information
Population Selected synthetic notes representing specified failure modes
Inputs Note text plus a versioned extraction policy
Outputs A strict JSON object with provenance spans
Reference Independently adjudicated field labels
Primary endpoint Exact field correctness under the declared schema
Critical failures Invented dose, wrong negation, wrong patient, unauthorized write
Comparator A simple rule-based extractor or previous configured system
Exclusions Predeclared corrupted files, not difficult cases discovered after scoring
Release rule Proposed thresholds approved for the actual use context

A threshold in a teaching example is an illustrative choice, not a clinical standard. In consequential deployments, set it with the accountable stakeholders, expected harm, comparator performance, and uncertainty.

Construct validity

An exam-style question may measure knowledge recall. A discharge summary eval may measure faithfulness, prioritization, and omissions. A tool-using trial matcher may measure data retrieval, eligibility logic, and permissions. These constructs overlap, but one result cannot stand in for all of them.

Write down what the eval deliberately does not measure. This prevents a later report from silently turning a narrow result into a broad capability claim.

Sampling and cases

Start with a development set covering normal cases and known failure mechanisms. Include easy controls to verify that the harness can succeed. Add difficult cases because they test a relevant risk, not because they make a model look bad. Create a separately frozen holdout for acceptance. Once you inspect it to improve prompts or graders, it is development data.

For population-level estimates, your sampling strategy must support the target population. A curated adversarial set supports statements about those adversarial cases, not prevalence of failures in routine use. Keep regression, stress, and representative cohorts separately named.

Counterfactual pairs

Construct pairs that differ in one relevant detail: a negation, missing date, permission state, or evidence version. A good system should change behavior when the detail matters and remain stable when an irrelevant detail changes. Counterfactuals test sensitivity to the actual rule rather than superficial vocabulary.

Exercise: create two trial-screening records identical except for an unknown eligibility variable. Under the supplied trial policy, the complete record can be “eligible” while the unknown record must be “needs review.” Explain why treating missing information as a negative value creates a false certainty.