Bryan Tegomoh / Learn
Chapter 21

From an eval notebook to a reliable evaluation service

TL;DR

Version cases, control execution, separate adapters from graders, and preserve complete reports. Expose failed and cancelled work rather than silently shrinking the denominator.

A notebook is useful for exploration. A dependable eval service needs stable inputs, controlled execution, complete run records, and honest failure states. Build only the infrastructure your current task requires.

Minimal architecture

Use five components: versioned cases, a runner, a controlled system adapter, graders, and a report store. Keep the reference-bearing dataset separate from model-visible inputs. The runner orchestrates work; the adapter performs the approved interaction; the grader judges artifacts; the report store preserves results.

Do not let the model modify its grader. Do not store the only copy of results in a chat session. Use content hashes and immutable run identifiers to identify inputs and outputs. A hash verifies identity of bytes, not truth of labels.

States and failure handling

Each run should be queued, running, completed, failed, or cancelled. Each case should record completion, timeout, invalid output, model error, tool error, or environment error. The report must expose incomplete work rather than silently presenting a smaller denominator.

Retry only under a declared rule. Preserve the original failure and each attempt. Keep the difference between retrying a transient provider failure and giving a model another chance to solve the task.

CI and regression

Continuous integration can run deterministic scorer tests and small offline regressions on each change. Paid model runs need separate budget and credential controls. Freeze model acceptance suites so a routine code change cannot edit away its failures.

After a scorer change, re-score saved outputs where valid. After a model or harness change, re-run relevant cases. These are different operations. Record which happened.

Observability

Track run IDs, case IDs, event timestamps, tool calls, tokens, costs, latency, resource exhaustion, and exceptions. Avoid collecting sensitive text merely because logging is convenient. Make it possible to find the first failing step from the evidence.

Operational monitoring complements offline evals. Watch real failure reports and distribution changes under the approved governance. A green regression suite can miss a new failure mechanism not represented in the suite.

Scaling decision

Before adding a queue or database, identify the bottleneck: parallel execution, durability, long-running work, auditability, or collaboration. Choose the smallest architecture that solves it. Document why a new dependency is needed and how failures will be detected.

Exercise: design an interrupted-run recovery procedure. It should identify unfinished cases, preserve completed artifacts, avoid duplicate side effects, and generate a report that clearly labels the resumed execution.