Bryan Tegomoh / Learn
Chapter 28

Lab walkthroughs with solutions and review notes

TL;DR

Work through repairs, strict scoring, missingness, rare-event metrics, temporal leakage, and judge checks. Solutions explain the mechanism and state exactly what each exercise proves.

Attempt the exercises before reading solutions when you want to test yourself. Reading a worked example is also useful, but label that as guided practice. The examples below use the downloadable kit and invented policies, not clinical guidelines or frontier-model observations.

Lab A: repair the label counter

From coding_lab, run python3 -m unittest -v. The baseline should fail whitespace normalization and empty-label handling while ordinary counting, input preservation, and empty-list behavior pass. This is an intentionally defective exercise implementation.

The smallest solution under the stated string-input contract is:

def count_labels(labels: list[str]) -> dict[str, int]:
    counts: dict[str, int] = {}
    for label in labels:
        key = label.strip().lower()
        if not key:
            raise ValueError("empty normalized label")
        counts[key] = counts.get(key, 0) + 1
    return counts

Why it works: normalization occurs before validation; invalid empty keys cannot enter the dictionary; the function iterates without assigning into the input list. Add an independent case mixing tabs, mixed case, and repeated labels. Do not add network access or a package for this task.

A reviewer should identify the two repaired behaviors, confirm existing behavior, inspect scope, and distinguish observed test success from untested runtime handling of non-string inputs. The contract accepts strings; a separate public-input parser would own validation of arbitrary JSON.

Lab B: strict scoring

Run the demo response file against the suite. The authored fixtures are designed to produce 15 exact matches out of 20 and a HOLD gate because selected critical cases fail. This expected result is determined by the fixture contents, not observed model ability.

Remove one response and rerun. The process must fail with missing-response evidence. Duplicate an ID and rerun: it must fail. Change a label to an unrecognized value: it must fail. Your scorer is defective if any of those operations quietly changes the denominator.

Export prompts and inspect the result. It must contain only IDs, prompts, and allowed labels. References and critical flags should be absent. Remember that export cannot protect the hidden suite if the agent can still read its folder.

Lab C: missing-data logic

Policy: absent fields stay unknown. Input: a medication is named; dose is not documented. A draft fills in a familiar dose. Correct judgment under this policy: review required because the draft invents missing information.

Counterfactual: dose is explicitly documented and the draft preserves it. Correct judgment: complete under the narrow extraction policy. Neither judgment says the regimen is medically appropriate.

Add a conflicting source entry. The correct output should surface the conflict rather than choose the convenient value silently. Ask whether your labels can represent this ambiguity and whether your report distinguishes it from ordinary missingness.

Lab D: rare-event arithmetic

Synthetic confusion matrix: TP 8, FN 2, FP 90, TN 900. Sensitivity = 8/10 = 0.8; specificity = 900/990, approximately 0.909; PPV = 8/98, approximately 0.082; accuracy = 908/1000 = 0.908.

The intended lesson is reviewer burden. Of 98 alerts, only 8 are true positives in this invented cohort. If review resources are scarce, accuracy alone is not the relevant operational summary. If missing an event is consequential, sensitivity and the missed-event cases still matter. Changing a threshold changes both workloads and errors.

Lab E: temporal leakage

A forecast is issued on day 10. A revised count is published on day 17. If you feed the revised count into the day-10 input, you have provided future information. Repair the backtest by storing and selecting the version available on day 10. Decide separately which target version you will score against.

Do not call a final-data retrospective forecast a prospective evaluation. The case can be useful for another question, but its label and interpretation must match the design.

Lab F: a model-judge stress test

Write a concise correct answer and a long incorrect answer under a supplied policy. Have a judge grade both without model identities. Reverse their order. Inspect whether its decision changes and whether its cited evidence supports the verdict. Save the outputs and record unresolved disagreements.

This is a small diagnostic exercise, not a complete judge-validation study. Expand with independently adjudicated cases before using the judge to gate a consequential system.