Bryan Tegomoh / Learn
Chapter 13

Diagnose failures and improve the system

TL;DR

Classify failures by mechanism so the next experiment has a purpose. Repair defective tasks or graders transparently and preserve a fresh acceptance set after tuning.

A useful failure analysis changes the next experiment. “The model hallucinated” is too broad to select a repair. Identify where the error entered the workflow and how you know.

An actionable taxonomy

Failure class Example Possible next experiment
Task ambiguity Reference assumes an unstated rule Clarify the specification and re-adjudicate
Knowledge failure Unsupported factual claim Supply a verified source and test use of evidence
Retrieval failure Correct source never returned Test query and retrieval coverage separately
Reasoning failure Applies exclusion rule backward Add counterfactual pairs and inspect decision logic
Tool failure Uses wrong argument or ignores denial Tighten interface and test error recovery
State failure Carries one patient's data to another Reset state and test isolation
Grader failure Rejects a valid alternative Repair grader against independently labeled controls
Environment failure Container fails before task starts Repair infrastructure and preserve aborted status
Policy failure Performs an unauthorized action Enforce tool permissions and test boundary cases

Keep multiple causes when the evidence supports them. Do not blame the model for a corrupted input, or blame the harness merely because the model failed.

The evidence packet

For each consequential failure retain task ID, input version, raw output, visible trace, environment revision, grader version, reference evidence, reproduction steps, and reviewer finding. Minimize sensitive content and enforce access controls. Public reports should include only publishable, safe evidence.

Fix one class at a time

Choose an intervention tied to the mechanism. A longer prompt may repair policy ambiguity but worsen context load. Retrieval may improve factual support but expose prompt injection. A stricter output schema may improve parsing while leaving semantics wrong. Tool enforcement can prevent unauthorized writes more reliably than a verbal reminder.

Run development cases and preserve regressions. Acceptance holdouts should answer whether the frozen candidate generalizes beyond the tuned cases. Avoid repeatedly peeking at the holdout until it becomes another development set.

When the benchmark is wrong

A failing test can reject a correct solution. Review the specification, test assumptions, and alternative implementation. Repair defective tasks transparently, version the suite, and recompute affected comparisons. Never quietly remove hard cases to improve a headline score.

OpenAI's coding-evaluation audit illustrates why benchmark artifacts themselves need scrutiny. This course does not recommend a public leaderboard as a substitute for validating your own tasks.

Exercise: an agent produces a correct API response with fields in a different order. Your grader rejects the text. Identify the grader defect, implement semantic JSON comparison, and add both a valid reordering and a wrong-value control.