Build datasets and reference standards
TL;DR
Build versioned cases with provenance and defensible labels. Keep answers out of model-visible inputs and separate development from a genuinely withheld acceptance set.
A dataset is not a pile of prompts. It is a versioned set of task instances with provenance, intended coverage, adjudication, and access rules. Weak datasets create convincing scores for the wrong question.
Start with an explicit schema
The starter kit uses small synthetic tasks with a common label interface:
{
"id": "med-001",
"track": "medicine",
"prompt": "Policy: absent fields remain unknown. Note: medication named, dose absent. Label?",
"labels": ["complete", "needs_review"],
"target": "needs_review",
"critical": true
}
This labeled suite is grader-side material. The export command produces model-visible prompts without target or critical. Never send the reference-bearing suite to the evaluated system. A benchmark that leaks its answer key is measuring access to answers.
python3 runner.py --suite suites.json --export-prompts prompts.json
The export is a convenience for this teaching kit, not a security boundary on a shared filesystem. A real agent environment needs access controls that prevent reading hidden graders and references.
Reference development
Write the labeling policy first. Have qualified annotators independently label a calibration subset. Capture disagreements before discussion. Resolve them with an adjudication rationale and source or policy evidence. Maintain ambiguous cases with an explicit unresolved status rather than forcing certainty.
Measure agreement appropriate to the label type and sampling design. Raw agreement can be high for imbalanced labels. Cohen's kappa has prevalence sensitivity and is not a universal quality certificate. Inspect disagreement categories, not only a summary statistic.
Provenance and licensing
Record origin, acquisition date, content version, permitted use, transformations, and any restrictions. Public access does not imply redistribution rights. Patient-level data requires a legitimate access and governance basis. Keep identifiers, private source text, and access credentials out of public eval artifacts.
Synthetic cases are useful for controlled failure mechanisms. They do not establish performance on real clinical records. Synthetic data can also inherit factual errors or stylistic regularities from the generator. Expert review and real-world external validation answer different questions.
Leakage and split design
Split at the level where related examples share information: patient, encounter, institution, study, outbreak, or repository. Near-duplicate prompts in both development and holdout can inflate performance. Future observations in a past forecast create temporal leakage even if filenames are different.
Use case IDs that do not reveal labels. Deduplicate content and related variants across splits. Document overlap with public benchmarks and prior training exposure when known. Unknown contamination is unknown, not proof of cleanliness.
Dataset review exercise
Take ten cases and inspect them manually. For each, identify the intended construct, reference evidence, alternative acceptable answer, and the smallest change that would alter the label. Remove or repair cases whose answer depends on an unstated assumption.
Acceptance criterion: two reviewers applying the same policy can explain the labels without seeing model responses, and each disputed label has a documented disposition.