Run your first complete evaluation
TL;DR
Run the local kit, inspect every score, and demonstrate rejection of invalid or incomplete responses. Fixture scores verify the apparatus, not a model or clinical workflow.
The downloadable kit demonstrates the complete measurement loop without calling a model. It includes a labeled suite, blinded prompt export, imperfect fixture responses, strict grading, a report, and failure-path tests. Its purpose is to make the mechanics visible.
Understand the files
| File | Purpose |
|---|---|
suites.json |
Synthetic grader-side cases and references |
prompts.json |
Exported model-visible cases, generated locally |
demo-responses.json |
Deliberately imperfect, authored response fixtures |
runner.py |
Validation, scoring, cohort summaries, and content hashes |
tests/test_runner.py |
Positive and negative scorer tests |
coding_lab/ |
A software-repair task with a documented bug |
templates/ |
Task contract, review rubric, and decision report |
The kit does not simulate hidden reasoning, call a frontier model, or estimate clinical accuracy. Running it proves that the scorer can process the supplied fixtures under the supplied policy.
Execute and inspect
python3 -m unittest discover -s tests -v
python3 runner.py --suite suites.json --export-prompts prompts.json
python3 runner.py --suite suites.json --responses demo-responses.json --output report.json
A successful process writes a complete report. The report can still say HOLD because evaluation execution and system acceptance are different outcomes. An invalid, duplicate, unknown, or missing response fails the process. Fix the data or investigate the harness; do not erase the offending case.
Replace fixtures with model observations
Provide only the blinded prompts to your approved system. Use the same declared system prompt, tool access, and budget for every case. Save outputs as a JSON list with exactly id and label per case. Preserve raw outputs separately, including invalid responses. A transformation into labels must have a declared rule and an audit trail.
Record model/provider identifier, date, prompt version, temperature where supported, output limit, tool settings, and retries in a separate run manifest. The starter report hashes inputs but does not supply this missing model provenance for you.
Use a provider-approved credential flow. Never put keys in the JSON files, source repository, browser JavaScript, or chat. Model calls can incur charges. Start only after a budget and permission are established. The site itself has no model API integration.
Read the report
Look at per-case correctness first, then track summaries and critical failures. The aggregate is a description of this cohort, not a generalized accuracy estimate. Small synthetic tracks have too few cases for meaningful deployment conclusions.
Choose one wrong case and inspect the prompt, raw output, expected label, and scoring rule. Ask whether the system, the reference, or the harness caused the discrepancy. Preserve that analysis before changing the prompt.
Your first improvement experiment
Freeze the existing fixture report as the baseline. If testing a real system, change one factor: a policy clarification, retrieval source, or tool schema. Re-run the same development cases and compare case-level changes. Check regressions as well as improvements. Test a fresh holdout only after tuning is complete.
Acceptance criterion: you can reproduce the report, explain every field, demonstrate a failure-path rejection, and state precisely what the result does not establish.