09. Review a coding agent end to end

Published

October 3, 2026

This is the closest track to the supplied G2i role. You review the interaction, repository changes, resulting behavior, and the quality of the agent’s technical judgment. Reading the final answer alone is insufficient.

NoteLearning objectives
  • Review the complete coding-agent trajectory.
  • Tie a finding to a violated requirement.
  • Reproduce a defect and verify the repair.
TipTL;DR

Review the full agent interaction, exact patch, resulting behavior, and verification. Tie findings to a violated requirement and reproducible evidence, not personal style preferences.

An evaluation handling sequence

  1. Read the task and environment constraints before watching the run.
  2. Identify expected behavior and plausible failure modes independently.
  3. Verify the initial repository and test baseline.
  4. Observe the transcript, tool calls, and scope changes.
  5. Inspect the final diff and files, including untracked artifacts.
  6. Run relevant checks against the final revision in a controlled environment.
  7. Separate task outcome from process quality.
  8. Write evidence-based findings and a bounded verdict.

A broken environment should be labeled an environment failure. An agent that cannot solve a valid task under the allowed resources has a capability failure. An ambiguous specification may require an invalid-task or unresolved classification. Do not force these into one “wrong answer” bucket.

Review rubric

Dimension Strong evidence Weak evidence
Task comprehension Reads constraints and identifies relevant files Starts rewriting unrelated modules
Debugging Reproduces failure and tests competing hypotheses Speculates and makes repeated blind edits
Correctness Covers positive, negative, and boundary behavior One happy-path demonstration
Architecture Fits existing interfaces and preserves invariants Adds abstractions without a use case
Scope Small coherent change tied to the task Unrequested dependency or migration
Verification Runs the right checks against the final code Says tests pass without execution
Communication Reports outcome, evidence, and limitations Overstates completion or hides failures
Security Respects permissions and protects sensitive artifacts Reads secrets or sends unauthorized data

Do not score an agent poorly merely because its approach differs from your preferred style. Judge whether the alternative is valid under the task contract. Reference patches are examples of solutions, not always the only acceptable implementation.

Worked trace: normalization bug

The task is to count a list of labels after trimming whitespace and converting to lowercase. An empty normalized label must raise an error. The input list must remain unchanged.

A weak agent changes label.lower() to label.strip().lower(), runs one example, and claims the issue is resolved. This handles surrounding whitespace but misses the empty-label rule. A stronger agent reads the full contract, reproduces both whitespace and empty-label failures, modifies the minimum function, and runs regression checks for ordinary labels and input preservation.

Both patches can look plausible. Only the second demonstrates the contract’s full behavior. You must inspect tests independently rather than rewarding the length of the agent’s explanation.

Write a surgical finding

Correctness: incomplete. The patch normalizes whitespace but accepts
an all-whitespace label as an empty dictionary key. The task requires
ValueError for that case. Reproduction: count_labels(["   "]).
The ordinary mixed-case example passes. Add the missing validation
and a regression test; retain the input-preservation check.

That is useful feedback because it names the violated requirement, evidence, and repair. “Not staff-level” without a concrete defect is not useful.

Exercise: complete the downloadable coding lab twice, first with a deliberately vague request and then with a task contract. Save both traces and compare observed behavior. Treat the comparison as a learning exercise, not a controlled model study unless the configurations were matched.