16. Medicine: clinical reasoning and documentation
Clinical evaluations require careful separation of tasks. Extracting a documented allergy, summarizing an encounter, estimating a diagnosis, and recommending treatment have different reference standards and risks. A correct answer on a vignette is not evidence of benefit in a clinical workflow.
- Preserve unknowns, negation, identity, and time.
- Separate extraction accuracy from treatment correctness.
- Bound the interpretation of software tests.
Separate extraction, documentation, clinical reasoning, and treatment tasks. Preserve unknowns, negation, entity identity, and time; software checks do not establish clinical benefit.
Worked task: faithful medication extraction
Use a supplied synthetic policy: report a medication’s name, status, dose, route, and frequency only when explicitly documented. Preserve unknown fields as unknown. Do not infer a dose from a familiar regimen. Record supporting spans and preserve negation and temporality.
{
"source": "Medication X was stopped last month. Dose not listed.",
"expected": {
"name": "Medication X",
"status": "stopped",
"dose": null,
"route": null,
"frequency": null
}
}Medication X is an invented placeholder, not a prescribing example. The task measures faithfulness to supplied text. It does not test whether stopping the medication was medically appropriate.
Build a reference standard
Create field-level labels and evidence spans. Have qualified reviewers adjudicate abbreviations, conflicting entries, and chronology under a written policy. Include “unable to determine” as a valid state when appropriate. Do not force an answer simply to keep every cell populated.
Grade each field, entity linkage, status, and support separately. Define which errors are critical for the actual use. A string match can check a name but may miss that it belongs to a different patient or time period.
Clinical reasoning evaluations
For reasoning tasks, use task-specific source evidence and a structured rubric: problem representation, differential relevance, evidence use, uncertainty, escalation, and contraindicated actions. Avoid making one panel’s preference the sole universal ground truth when multiple defensible decisions exist.
Include insufficient-information cases. Evaluate whether the system requests the missing information or communicates the limitation. Test inappropriate reassurance and unsupported specificity, not only overt factual errors.
Evidence ladder
Start with retrospective task performance and error analysis. Then consider external validation, clinician-in-the-loop testing, and prospective workflow impact under the appropriate governance. Different evidence addresses different claims. Do not describe software test success as clinical validation.
The TRIPOD-LLM article is a reporting guideline for studies using LLMs. Use the paper and its checklist to improve study reporting, not as a certificate that a model is clinically safe. The article link was retrieved in search; direct full-text access was unavailable during course preparation.
Starter-kit connection
The medicine cases test whether missing or contradictory information triggers review under a supplied toy policy. They are not clinical triage guidance. The meaningful next project is a small expert-adjudicated extraction cohort with transparent unknown handling.
Exercise: create a note with one historical medication and one current medication. Design a grader that catches swapped statuses, invented doses, and loss of negation. Explain why a high text-similarity score could miss all three.