Bryan Tegomoh / Learn
Chapter 16

Medicine: clinical reasoning and documentation

TL;DR

Separate extraction, documentation, clinical reasoning, and treatment tasks. Preserve unknowns, negation, entity identity, and time; software checks do not establish clinical benefit.

Clinical evaluations require careful separation of tasks. Extracting a documented allergy, summarizing an encounter, estimating a diagnosis, and recommending treatment have different reference standards and risks. A correct answer on a vignette is not evidence of benefit in a clinical workflow.

Worked task: faithful medication extraction

Use a supplied synthetic policy: report a medication's name, status, dose, route, and frequency only when explicitly documented. Preserve unknown fields as unknown. Do not infer a dose from a familiar regimen. Record supporting spans and preserve negation and temporality.

{
  "source": "Medication X was stopped last month. Dose not listed.",
  "expected": {
    "name": "Medication X",
    "status": "stopped",
    "dose": null,
    "route": null,
    "frequency": null
  }
}

Medication X is an invented placeholder, not a prescribing example. The task measures faithfulness to supplied text. It does not test whether stopping the medication was medically appropriate.

Build a reference standard

Create field-level labels and evidence spans. Have qualified reviewers adjudicate abbreviations, conflicting entries, and chronology under a written policy. Include “unable to determine” as a valid state when appropriate. Do not force an answer simply to keep every cell populated.

Grade each field, entity linkage, status, and support separately. Define which errors are critical for the actual use. A string match can check a name but may miss that it belongs to a different patient or time period.

Clinical reasoning evaluations

For reasoning tasks, use task-specific source evidence and a structured rubric: problem representation, differential relevance, evidence use, uncertainty, escalation, and contraindicated actions. Avoid making one panel's preference the sole universal ground truth when multiple defensible decisions exist.

Include insufficient-information cases. Evaluate whether the system requests the missing information or communicates the limitation. Test inappropriate reassurance and unsupported specificity, not only overt factual errors.

Evidence ladder

Start with retrospective task performance and error analysis. Then consider external validation, clinician-in-the-loop testing, and prospective workflow impact under the appropriate governance. Different evidence addresses different claims. Do not describe software test success as clinical validation.

The TRIPOD-LLM article is a reporting guideline for studies using LLMs. Use the paper and its checklist to improve study reporting, not as a certificate that a model is clinically safe. The article link was retrieved in search; direct full-text access was unavailable during course preparation.

Starter-kit connection

The medicine cases test whether missing or contradictory information triggers review under a supplied toy policy. They are not clinical triage guidance. The meaningful next project is a small expert-adjudicated extraction cohort with transparent unknown handling.

Exercise: create a note with one historical medication and one current medication. Design a grader that catches swapped statuses, invented doses, and loss of negation. Explain why a high text-similarity score could miss all three.