14. Write reviews, calibration notes, and decision reports

Published

October 3, 2026

A review should let another qualified person inspect the evidence and reach a judgment. Concise writing is not shorthand for vague writing. State the result, the specific observation, and the consequence.

NoteLearning objectives
  • Write findings tied to requirements and evidence.
  • State consequences and unresolved uncertainty.
  • Recommend a decision the experiment can support.
TipTL;DR

Write the result first, then the requirement, evidence, consequence, and uncertainty. A good report supports a specific decision without overstating what was tested.

A review record

Task: normalized-label counter, revision [record actual revision]
Outcome: incomplete implementation
Finding: whitespace-only labels become an empty key instead of raising
Evidence: count_labels(["   "]) returns {"": 1}; contract requires ValueError
Scope: ordinary case normalization works; no evidence of input mutation
Process: agent ran only the happy-path example, not the specified suite
Recommendation: repair validation and rerun the complete relevant checks
Uncertainty: no performance or concurrency requirement was tested

Bracketed fields in templates are instructions to supply observed values. They are not fabricated examples of execution.

Pairwise preference

When comparing two responses, identify the decisive criterion before ranking. If A is concise but wrong and B is correct with modest extra explanation, correctness may govern. If both are correct, compare completeness, scope, evidence, and usability under the task contract.

Do not prefer verbose responses by default. Do not penalize a response for acknowledging uncertainty when the input does not support certainty. Distinguish calibrated uncertainty from evasion.

Calibration meetings

Have reviewers independently judge the same cases. Discuss disagreements using concrete evidence. Update rubric anchors when disagreement reflects ambiguity, not merely reviewer taste. Keep the original ratings so agreement before and after calibration can be assessed.

If the rubric changes, decide whether prior cases must be re-rated. A score under rubric v1 is not directly comparable with rubric v2 unless the mapping is justified.

A decision-ready report

Lead with the action: advance to a limited pilot, hold, revise the experiment, or reject the tested use. Include task and population, configured system, cohort, comparator, primary outcome, uncertainty, critical failures, subgroup findings, resource use, exclusions, and limitations. Link each consequential claim to evidence.

Separate executed checks from proposed work. “All unit tests passed” is software verification. “The system is safe for patients” requires a different and much stronger body of evidence. Do not allow formatting to collapse those categories.

Staff-level communication

Staff-level judgment often appears in what a reviewer does not overclaim. It names tradeoffs, identifies the bottleneck, proposes the smallest meaningful repair, and anticipates how a change affects other parts of the system. It also identifies when evidence is insufficient to make the decision.

Exercise: rewrite “The agent did a great job and the code looks clean” into two concrete findings, one successful behavior and one limitation, with reproduction evidence.