Biosurveillance and epidemic intelligence evaluations
TL;DR
Preserve information available at the decision time and define event-level targets. Evaluate delay, alert burden, uncertainty, and revisions instead of relying on a single retrospective score.
Biosurveillance evals must account for time, delayed observations, revisions, and decision context. A system can predict a revised historical series well while failing under the information available at the original decision time.
Define the task precisely
Distinguish signal detection, report summarization, nowcasting, forecasting, and prioritization. A summarizer extracts supplied evidence. A detector identifies a possible unusual event. A forecaster predicts a future target. They need different references and metrics.
For event detection, define the event, geography, observation window, evidence threshold, and escalation policy. A rumor, an unverified report, and a confirmed event are different evidence states. The model must not collapse them.
Worked task: review queue
Use a toy policy: a complete, corroborated report enters the review queue; an incomplete or contradictory report requires further verification; a duplicate is linked to the existing event rather than treated as a new one. The labels are workflow states, not public health declarations.
Create synthetic reports with changing place names, dates, duplicate content, missing denominators, and source corrections. Grade preservation of evidence state, event identity, uncertainty, and appropriate routing.
Temporal backtesting
For each historical decision date, reconstruct only information available then. Store publication timestamps, ingestion timestamps, and revision versions. Do not use final revised data as input to a retrospective forecast issued earlier. Evaluate against a declared target version and explain revisions.
Use rolling-origin evaluation. Fit or configure on earlier windows, predict the next window, advance, and repeat. Preserve outbreak or region clusters when assessing uncertainty. Avoid mixing overlapping targets as if they were independent observations.
Forecast metrics
MAE summarizes absolute point error. Interval coverage asks how often the declared interval contains the observation. Coverage alone can reward very wide intervals, so also assess sharpness and a proper scoring rule appropriate to the forecast format.
For a central interval [l, u] with nominal coverage 1-alpha, the interval score is (u-l) + (2/alpha)*(l-y) when y<l, and (u-l) + (2/alpha)*(y-u) when y>u; inside the interval it is just width. Lower is better. Weighted interval score combines interval scores and the median error with declared weights.
The epidemic-forecast scoring paper is a primary methodological reference for WIS. The linked version is a preprint. Use its definitions when implementing the metric, and test against an independent reference implementation before reporting research results.
Operational metrics
Track detection delay, event-level sensitivity, false alerts per unit time, review burden, geographic coverage, and missed-event severity. Do not define every negative report as an independent true negative if many concern the same event.
Exercise: construct a synthetic event whose case count is revised a week later. Write two manifests showing the original and revised information. Explain which version belongs in the forecast input and which target definition you will use for scoring.