08. Run your first complete evaluation

Published

October 3, 2026

The downloadable kit demonstrates the complete measurement loop without calling a model. It includes a labeled suite, blinded prompt export, imperfect fixture responses, strict grading, a report, and failure-path tests. Its purpose is to make the mechanics visible.

NoteLearning objectives
  • Run the complete local evaluation kit.
  • Export prompts without hidden references.
  • Demonstrate rejection of incomplete or invalid responses.
TipTL;DR

Run the local kit, inspect every score, and demonstrate rejection of invalid or incomplete responses. Fixture scores verify the apparatus, not a model or clinical workflow.

Understand the files

File Purpose
suites.json Synthetic grader-side cases and references
prompts.json Exported model-visible cases, generated locally
demo-responses.json Deliberately imperfect, authored response fixtures
runner.py Validation, scoring, cohort summaries, and content hashes
tests/test_runner.py Positive and negative scorer tests
coding_lab/ A software-repair task with a documented bug
templates/ Task contract, review rubric, and decision report

The kit does not simulate hidden reasoning, call a frontier model, or estimate clinical accuracy. Running it proves that the scorer can process the supplied fixtures under the supplied policy.

Execute and inspect

python3 -m unittest discover -s tests -v
python3 runner.py --suite suites.json --export-prompts prompts.json
python3 runner.py --suite suites.json --responses demo-responses.json --output report.json

A successful process writes a complete report. The report can still say HOLD because evaluation execution and system acceptance are different outcomes. An invalid, duplicate, unknown, or missing response fails the process. Fix the data or investigate the harness; do not erase the offending case.

Replace fixtures with model observations

Provide only the blinded prompts to your approved system. Use the same declared system prompt, tool access, and budget for every case. Save outputs as a JSON list with exactly id and label per case. Preserve raw outputs separately, including invalid responses. A transformation into labels must have a declared rule and an audit trail.

Record model/provider identifier, date, prompt version, temperature where supported, output limit, tool settings, and retries in a separate run manifest. The starter report hashes inputs but does not supply this missing model provenance for you.

Use a provider-approved credential flow. Never put keys in the JSON files, source repository, browser JavaScript, or chat. Model calls can incur charges. Start only after a budget and permission are established. The site itself has no model API integration.

Read the report

Look at per-case correctness first, then track summaries and critical failures. The aggregate is a description of this cohort, not a generalized accuracy estimate. Small synthetic tracks have too few cases for meaningful deployment conclusions.

Choose one wrong case and inspect the prompt, raw output, expected label, and scoring rule. Ask whether the system, the reference, or the harness caused the discrepancy. Preserve that analysis before changing the prompt.

Your first improvement experiment

Freeze the existing fixture report as the baseline. If testing a real system, change one factor: a policy clarification, retrieval source, or tool schema. Re-run the same development cases and compare case-level changes. Check regressions as well as improvements. Test a fresh holdout only after tuning is complete.

Acceptance criterion: you can reproduce the report, explain every field, demonstrate a failure-path rejection, and state precisely what the result does not establish.