10. Debug systematically and assess architecture

Published

October 3, 2026

Debugging is hypothesis testing over a program. Your epidemiology training helps here: identify the observed failure, define competing explanations, find discriminating observations, and avoid changing many variables at once.

NoteLearning objectives
  • Reproduce failures before proposing a fix.
  • Compare causal hypotheses with controlled tests.
  • Assess architecture against state and failure requirements.
TipTL;DR

Reproduce the failure, test competing hypotheses, and repair the smallest verified cause. Judge architecture by state ownership, failure handling, and actual requirements.

Reproduce before repairing

Write the smallest input that produces the failure. Capture actual versus expected behavior, error output, environment, and revision. Repeat it to distinguish persistent from intermittent behavior. If the issue cannot be reproduced, preserve that uncertainty rather than fabricating a confident cause.

Trace the input through validation, transformation, storage, and presentation. For a wrong model score, the cause may be label normalization, missing case alignment, a stale reference, or an incorrect model output. Debugging the model before checking the scorer wastes effort.

A hypothesis table

Observation Hypothesis Discriminating check
Report omits a difficult case Parser silently skipped invalid JSON Submit one malformed record and inspect exit behavior
Different scores on rerun Shared environment retained state Reset to the same snapshot and compare
Retrieval answer has wrong date Old source or wrong source selected Inspect retrieved document version and evidence span
Test passes locally, fails in CI Dependency or platform difference Compare environment manifests and lockfiles

Do not equate correlation with cause. A failure disappearing after a dependency update does not prove which change repaired it. Reduce the explanation to the narrowest verified cause.

Architecture questions an evaluator must ask

What is the source of truth? Where are inputs validated? Which component owns state? What happens when a tool fails halfway through? How are retries deduplicated? What data is cached, and when does it expire? Can different patients or tenants see each other’s information? What is observable when the process hangs?

A small local scorer does not need a distributed queue. A production evaluation service handling long-running agent tasks may need a queue, cancellation, resource isolation, and durable run records. The correct architecture follows requirements and failure modes, not fashionable tools.

Reliability tradeoffs

Retries improve recovery from transient errors but can duplicate side effects and selectively alter the measured cohort. Timeouts bound execution but can penalize legitimate long tasks. Caching saves cost but can invalidate a test of current retrieval. Concurrency increases throughput but introduces rate limits and resource contention.

Record the chosen tradeoff in the eval configuration. If model A gets more retries, longer timeouts, or stronger tools than model B, the comparison is between systems with different resources. That can be useful, but say so.

Exercise: interrupted report write

An eval finishes scoring, but the process is interrupted while writing the JSON report. Should a later reader accept the partial file? No. Design atomic output publication: write a temporary complete file, flush it, and replace the destination only when valid. Preserve a run status that distinguishes started, failed, and completed work.

Acceptance criterion: the final design specifies behavior for normal completion, invalid input, timeout, interruption, and repeated execution.