Health systems and public health evaluations
TL;DR
Distinguish administrative, operational, population, and patient-facing tasks. Evaluate routing and workload as well as content, with explicit references and governance.
Health AI includes administrative, operational, population, and patient-facing systems. The first task is to distinguish which of those you are evaluating. A hospital scheduling classifier and an autonomous treatment recommender cannot share a release policy merely because both involve healthcare.
Worked task: referral-document completeness
Use a synthetic policy: a referral packet requires a reason, a destination, and an authorization status. Missing or contradictory fields require manual review. The task is administrative completeness, not clinical appropriateness.
Input: “Reason documented; destination documented; authorization not recorded.” Expected label: needs_review. The system should not infer authorization from the presence of the referral. Create counterfactuals with explicit authorization, contradiction, and irrelevant details.
A deterministic grader can compare the label. For a richer extraction output, require field values and supporting spans. A human adjudicator should assess whether the policy itself captures the intended workflow.
Evaluation design
Stratify by organization, document type, language, and data availability when those affect intended use. Split related packets at patient or referral level. Measure review workload, incorrect auto-completion, missing critical fields, and user correction rates. An overall score can hide poor performance in a less common document type.
Test the workflow around the classifier: routing, queue delays, duplicate referrals, stale authorization, and handoff to a human. A correct label that never reaches the intended reviewer is an operational failure.
Population health example
Evaluate a system summarizing surveillance reports. Define source-bounded claims: geography, observation period, case definition, denominator, uncertainty, and data limitations. Penalize incorrect denominators and unsupported causal conclusions separately from missing stylistic details.
Ask whether the summary preserves the distinction between reported counts and underlying incidence. A change in reporting or testing can alter observed counts. The task should require appropriate qualification when the supplied source identifies that limitation.
Human factors
Measure how reviewers use the output. Does a high-confidence visual label discourage checking? Can they recover the original evidence? Does the system provide a correction route? A retrospective output study does not establish workflow benefit; a prospective supervised study tests a different question.
Starter-kit connection
The health track contains synthetic completeness cases. Use them to test the label pipeline and practice a report. Expand into a local workflow only after defining governance, a reference standard, and data access.
Exercise: write an eval in which an administrative output is technically correct but routed to the wrong organization. Define separate scores for content and routing, and identify the critical failure.