05. Design an evaluation before building it
An eval should test an explicit claim. “This model is good at medicine” is too broad. “Under the supplied synthetic reconciliation policy, this configured system preserves medication status and does not invent missing dose fields” is testable.
- Specify task, population, reference, and consequence.
- Distinguish the measured construct from broader claims.
- Write a falsifiable acceptance rule.
Define the decision, task, population, inputs, outputs, reference, and failure criteria before building. Match the strength of your conclusion to the construct actually tested.
The evaluation contract
Use the downloadable task-contract template. Complete every field before implementation:
| Field | Example |
|---|---|
| Decision | Whether to advance an extraction assistant to a supervised pilot |
| Intended user | A clinician reviewing a draft, not an autonomous prescribing system |
| Unit | One encounter note and its structured output |
| Construct | Faithful extraction of stated information |
| Population | Selected synthetic notes representing specified failure modes |
| Inputs | Note text plus a versioned extraction policy |
| Outputs | A strict JSON object with provenance spans |
| Reference | Independently adjudicated field labels |
| Primary endpoint | Exact field correctness under the declared schema |
| Critical failures | Invented dose, wrong negation, wrong patient, unauthorized write |
| Comparator | A simple rule-based extractor or previous configured system |
| Exclusions | Predeclared corrupted files, not difficult cases discovered after scoring |
| Release rule | Proposed thresholds approved for the actual use context |
A threshold in a teaching example is an illustrative choice, not a clinical standard. In consequential deployments, set it with the accountable stakeholders, expected harm, comparator performance, and uncertainty.
Construct validity
An exam-style question may measure knowledge recall. A discharge summary eval may measure faithfulness, prioritization, and omissions. A tool-using trial matcher may measure data retrieval, eligibility logic, and permissions. These constructs overlap, but one result cannot stand in for all of them.
Write down what the eval deliberately does not measure. This prevents a later report from silently turning a narrow result into a broad capability claim.
Sampling and cases
Start with a development set covering normal cases and known failure mechanisms. Include easy controls to verify that the harness can succeed. Add difficult cases because they test a relevant risk, not because they make a model look bad. Create a separately frozen holdout for acceptance. Once you inspect it to improve prompts or graders, it is development data.
For population-level estimates, your sampling strategy must support the target population. A curated adversarial set supports statements about those adversarial cases, not prevalence of failures in routine use. Keep regression, stress, and representative cohorts separately named.
Counterfactual pairs
Construct pairs that differ in one relevant detail: a negation, missing date, permission state, or evidence version. A good system should change behavior when the detail matters and remain stable when an irrelevant detail changes. Counterfactuals test sensitivity to the actual rule rather than superficial vocabulary.
Exercise: create two trial-screening records identical except for an unknown eligibility variable. Under the supplied trial policy, the complete record can be “eligible” while the unknown record must be “needs review.” Explain why treating missing information as a negative value creates a false certainty.