41. Build a biosurveillance benchmark end to end

Published

October 3, 2026

A biosurveillance benchmark specifies a decision, reconstructs the information available when that decision was made, and scores the resulting workflow against an independently adjudicated reference. Reports are observations; events are the units the workflow cares about. A useful benchmark preserves that distinction through ingestion, retrieval, reasoning, queue routing, and evaluation.

NoteLearning objectives
  • Replay only report revisions available at the decision time.
  • Score events, unsupported episodes, delays, and queue burden separately.
  • Reject temporal leakage and invalid benchmark references.
TipTL;DR

Build a temporally faithful event-review benchmark before adding an LLM. The runnable fixture separates report revisions from hidden event references and exposes unsupported escalations even when recall is unchanged. Preserve event-level denominators and queue burden.

Introduction

Lesson 18 owns the temporal and forecasting concepts. This exercise implements a small event-review benchmark, not an epidemic forecast or a public-alert policy. The data are invented report IDs and evidence states with no patient data, pathogen procedures, or live sources.

Write the benchmark card before implementation

Decision: compare evidence-preserving review-queue candidates.
Task: propose one review item per report when corroboration is available.
Input: only report revisions available by the replay day.
Output: report ID and proposed review day.
Reference: grader-only event linkage and evidence-readiness day.
Baseline: deterministic corroboration rule, no LLM.
Candidate: deliberately premature rule used to test the grader.
Primary dimensions: detected events, unsupported episodes, detection delay.
Operational dimensions: raw review items and duplicate-event alerts.
Window: six invented days; all reference events fall within the window.
Claim limit: operational scoring demonstration, not model or field validation.

The reference readiness day is the first day the supplied benchmark policy treats an event as actionable. It is not symptom onset, infection time, or retrospectively discovered outbreak onset. In a real study, name and justify each time target separately.

What the data must contain

Artifact Required fields or decisions Why it matters
Source report Stable ID, source identity, publication and ingestion time Reconstruct what was actually available
Revision Revision ID, availability time, correction or withdrawal Prevent later information from leaking backward
Evidence annotation Supporting spans, evidence state, ambiguity Distinguish a claim from corroboration
Event ledger Event ID, geography, time window, inclusion rule Separate one event from many reports
Linkage reference Report-to-event mapping and adjudication Grade duplicates without exposing the answer
Replay manifest Cutoffs, source versions, configuration, failures Reproduce a decision sequence
Output artifact Action, evidence IDs, uncertainty, proposed time Score grounded behavior and routing
Operational reference Queue ownership, deadline, severity policy Evaluate consequences and workload

This teaching program reduces publication and ingestion to one integer availability day. It treats evidence state as supplied metadata and linkage as a grader reference. Real extraction, corroboration, multilingual handling, and event linking need their own evaluated components.

Run the complete replay

From the downloaded eval-starter-kit folder:

python3 foundations/biosurveillance_benchmark.py
python3 -m unittest discover -s tests -v
from __future__ import annotations

import json
from dataclasses import asdict, dataclass


@dataclass(frozen=True)
class Report:
    report_id: str
    revision: int
    available_day: int
    evidence: str


@dataclass(frozen=True)
class Alert:
    report_id: str
    day: int


def snapshot(reports: list[Report], day: int) -> dict[str, Report]:
    if type(day) is not int or day < 1:
        raise ValueError("day must be a positive integer")
    seen: set[tuple[str, int]] = set()
    revisions: dict[str, list[Report]] = {}
    for r in reports:
        if (not r.report_id.strip() or type(r.revision) is not int or r.revision < 1
                or type(r.available_day) is not int or r.available_day < 1
                or r.evidence not in {"unverified", "corroborated", "withdrawn"}):
            raise ValueError("invalid report")
        key = (r.report_id, r.revision)
        if key in seen:
            raise ValueError("duplicate report revision")
        seen.add(key)
        revisions.setdefault(r.report_id, []).append(r)
    current: dict[str, Report] = {}
    for report_id, versions in revisions.items():
        versions.sort(key=lambda r: r.revision)
        if any(a.available_day > b.available_day for a, b in zip(versions, versions[1:])):
            raise ValueError("revision availability is out of order")
        eligible = [r for r in versions if r.available_day <= day]
        if eligible:
            current[report_id] = eligible[-1]
    return current


def replay(reports: list[Report], days: int, require_corroboration: bool) -> list[Alert]:
    if type(days) is not int or days < 1 or type(require_corroboration) is not bool:
        raise ValueError("invalid replay configuration")
    emitted: set[str] = set()
    alerts: list[Alert] = []
    for day in range(1, days + 1):
        for r in snapshot(reports, day).values():
            eligible = r.evidence == "corroborated" if require_corroboration else r.evidence != "withdrawn"
            if eligible and r.report_id not in emitted:
                alerts.append(Alert(r.report_id, day))
                emitted.add(r.report_id)
    return alerts


def score(reports: list[Report], alerts: list[Alert], event_map: dict[str, str | None],
          event_ready: dict[str, int], days: int) -> dict[str, object]:
    if type(days) is not int or days < 1 or not event_ready:
        raise ValueError("positive window and nonempty event reference required")
    if set(event_map) != {r.report_id for r in reports}:
        raise ValueError("reference map must cover exactly the report IDs")
    if any(not k.strip() or type(v) is not int or not 1 <= v <= days for k, v in event_ready.items()):
        raise ValueError("invalid reference readiness day")
    if any(e is not None and e not in event_ready for e in event_map.values()):
        raise ValueError("unknown reference event")
    detected: dict[str, int] = {}
    unsupported: set[str] = set()
    duplicates = 0
    seen: set[str] = set()
    for alert in sorted(alerts, key=lambda a: a.day):
        if type(alert.day) is not int or not 1 <= alert.day <= days or alert.report_id in seen:
            raise ValueError("invalid or repeated alert")
        seen.add(alert.report_id)
        current = snapshot(reports, alert.day)
        if alert.report_id not in current:
            raise ValueError("alert cites absent or future information")
        r = current[alert.report_id]
        event = event_map[alert.report_id]
        if r.evidence != "corroborated" or event is None or alert.day < event_ready[event]:
            unsupported.add(alert.report_id)
        elif event in detected:
            duplicates += 1
        else:
            detected[event] = alert.day
    delays = {e: day - event_ready[e] for e, day in detected.items()}
    episodes = len(detected) + len(unsupported)
    return {"fixture_kind": "synthetic operational backtest, not an LLM evaluation",
            "days": days, "reference_events": len(event_ready), "detected_events": len(detected),
            "event_recall": len(detected) / len(event_ready),
            "episode_precision": len(detected) / episodes if episodes else None,
            "unsupported_episodes": len(unsupported), "unsupported_per_day": len(unsupported) / days,
            "duplicate_event_alerts": duplicates, "raw_review_items": len(alerts),
            "detection_delay_days": delays,
            "missed_events": sorted(set(event_ready) - set(detected))}


def fixtures() -> tuple[list[Report], dict[str, str | None], dict[str, int]]:
    reports = [Report("r-a", 1, 2, "unverified"), Report("r-a", 2, 3, "corroborated"),
               Report("r-a-copy", 1, 4, "corroborated"), Report("r-b", 1, 5, "corroborated"),
               Report("r-rumor", 1, 2, "unverified"), Report("r-rumor", 2, 6, "withdrawn")]
    # Grader-only event linkage is never passed to replay().
    return reports, {"r-a": "event-a", "r-a-copy": "event-a", "r-b": "event-b", "r-rumor": None}, {"event-a": 3, "event-b": 5}


def main() -> None:
    reports, event_map, event_ready = fixtures()
    for name, strict in [("corroboration_baseline", True), ("premature_candidate", False)]:
        alerts = replay(reports, 6, strict)
        print(json.dumps({"system": name, "alerts": [asdict(a) for a in alerts],
                          **score(reports, alerts, event_map, event_ready, 6)}, sort_keys=True))


if __name__ == "__main__":
    main()

snapshot selects the latest eligible revision for each report. replay has no access to the hidden event mapping. score resolves emitted report IDs to event references, checks evidence eligibility, and computes the workflow outcomes. A candidate cannot score itself by returning an event ID copied from the reference.

The fixture has an unverified report later corroborated, a syndicated copy linked to the same event, a second event, and a rumor later withdrawn. Withdrawal on day six must not appear in the day-two snapshot. The code rejects duplicated revisions and revisions whose availability goes backward.

Read the observed fixture comparison

Outcome Corroboration baseline Premature candidate
Reference events 2 2
Detected events 2 2
Event recall 1.0 1.0
Episode precision under the teaching policy 1.0 0.5
Unsupported episodes 0 2
Raw review items 3 4
Duplicate valid event alerts 1 0
Event A valid detection delay 0 days 1 day
Event B valid detection delay 0 days 0 days

These results were produced by the supplied deterministic programs. They are not observations of an LLM or estimates of surveillance performance. The two-event fixture is designed to expose scoring mechanisms, not support population inference.

The premature candidate emits r-a while unverified. That is an unsupported escalation under the teaching rule, even though the underlying event is later corroborated. It emits that report only once; its first valid event-A detection therefore comes from the syndicated copy on day four. This explains the one-day valid detection delay.

The baseline detects both events at their reference readiness times, but still creates a duplicate item for the syndicated copy. Improving evidence state is not the same as solving deduplication. The zero duplicate count for the premature candidate also looks attractive in isolation, but it partly reflects an earlier unsupported item. Read the dimensions together.

Define the denominators explicitly

Event recall is detected reference events divided by eligible reference events. Valid repeated alerts for the same event do not create additional true-positive events. Detection delay is measured only for detected events; missed events remain listed separately instead of disappearing from the report.

For this fixture, episode precision uses unique detected valid events plus unsupported report episodes in its denominator. Each report emits at most once. In a real benchmark, false-event clustering must be adjudicated too: many reports of the same false rumor should not automatically become independent false events.

Raw review items measure queue burden before deduplication. Unsupported episodes per day use the declared observation window. The code does not manufacture a true-negative count from every document that did not generate an alert. Consequently, it does not report specificity for this event-review design.

If no episode is emitted, precision is null, not a falsely reassuring perfect score. The event reference must still be nonempty. A different benchmark may deliberately include event-free windows; predeclare how such windows contribute to workload and false-alert reporting.

Failure paths that validate the benchmark

Try an alert citing a report before it was available. The scorer raises. Add a second revision with an earlier availability day than its predecessor. Snapshot reconstruction raises. Repeat an alert for the same report, omit a reference mapping, use an unknown event, or put an alert outside the observation window. Each is rejected.

Those checks stop implementation errors from masquerading as performance. They do not validate the domain reference. Qualified reviewers must adjudicate whether the events, evidence states, matching windows, and operational policy describe the intended use.

Add an LLM without changing the question

Keep the input snapshot and output contract fixed. Ask the model to preserve the supplied evidence state and propose an eligible review item with supporting report IDs. Save the full response and parse it strictly. Represent needs_verification, no action, provider error, and missing output explicitly under the predeclared run policy.

First compare a simple model call with the deterministic baseline. Add retrieval only when the task requires finding evidence; evaluate eligible-source retrieval separately. Add tools only when an authorized workflow requires them. The agent loop in lesson 38 supplies the execution pattern.

If the model performs event linking, remove grader-only event IDs from its visible inputs. Evaluate known duplicates and near-duplicates representing genuinely different events. Keep an adjudicated mapping for scoring. A perfectly deduplicated queue that merges distinct events is a failure.

Expand into a serious research benchmark

Select sources under their actual access terms and your organization’s data policy. Preserve provenance and ingestion failures. Sample event-free periods, different reporting intensities, languages, locations, delayed corrections, and changing source coverage. Avoid convenience cases selected only because one system succeeds on them.

Split by event and time; inspect near-duplicate articles across partitions. Freeze an independent acceptance set. Use blinded domain review, written matching criteria, disagreement records, and appropriate event-level uncertainty. Compare fixed system configurations on matched windows and preserve missing runs.

Evaluate evidence extraction, retrieval, linking, routing, and end-to-end results separately. Measure reviewer burden and escalation timeliness under supervised workflow conditions when that is the intended claim. Forecasting is a separate track with target definitions and scoring rules in lesson 18.

The release decision requires more than benchmark performance: authorized access, tested containment, operational ownership, recovery behavior, and evidence that the workflow is appropriate. A personal learning exercise can teach that structure without pretending to supply field validation.

A concrete completion artifact

Recommendation: retain the corroboration baseline for the teaching comparison.
Evidence: both systems detect two events; candidate has two unsupported episodes.
Residual defect: baseline creates one valid duplicate-event review item.
Next experiment: event linking under the same eligible-source replay.
Protected behavior: distinct events must not be merged; future revisions stay excluded.
Limit: invented metadata fixture, no measured LLM or deployment outcome.

Checkpoint: add a third event with delayed reporting and a false rumor with two syndicated copies in a local copy. Write the reference first, predict the denominators, run the scorer, and explain what needs to change to cluster false episodes correctly.