12. Statistics, uncertainty, and comparisons
An eval score is an estimate or a cohort description, depending on design. The denominator, unit of analysis, sampling, and dependence structure determine what it means. Treat those as part of the result.
- Name denominators and dependence structures.
- Interpret uncertainty in matched comparisons.
- Keep severe errors visible beside averages.
Name the denominator, sampling unit, dependence, and uncertainty before interpreting a score. Compare matched cases and keep critical failures visible beside averages.
Metrics with explicit denominators
Sensitivity = TP / (TP + FN). Specificity = TN / (TN + FP). Positive predictive value = TP / (TP + FP). Accuracy = (TP + TN) / N. F1 = 2TP / (2TP + FP + FN). A zero denominator produces an undefined metric, not a zero score.
For rare outcomes, a high accuracy can coexist with poor detection. Consider this synthetic arithmetic example: TP = 8, FN = 2, FP = 90, TN = 900. Sensitivity is 80%, while PPV is approximately 8.2%. The alert queue contains many more false alerts than true alerts. These numbers are invented for teaching, not measured system performance.
Use the site’s interactive calculator to explore changes in prevalence and false positives. Ask which metric matters to the actual workflow: missed urgent cases, reviewer workload, or correctly excluded routine cases.
Confidence intervals
The kit reports a Wilson interval for the pooled binary correctness proportion. That is a descriptive calculation under an independent Bernoulli interpretation, not proof that the synthetic case collection represents a target population. Correlated cases, repeated patient encounters, or curated failure families need different uncertainty treatment.
Repeated trials on one case are not independent patients. Report within-case variability separately. For clustered sampling, bootstrap at the cluster level when justified. Explain assumptions and the limitations of small cohorts.
Zero observed failures
Zero observed failures does not establish zero risk. Under independent identically distributed Bernoulli trials, the one-sided exact 95% upper bound after zero failures in n trials is 1 - 0.05**(1/n). The approximate rule of three gives 3/n for sufficiently large n. Neither applies automatically to adversarial or clustered samples.
If you want an upper bound below a proposed risk tolerance, derive the sample size from the tolerance, confidence, and assumptions. Do not choose the sample size by how many prompts happen to be convenient.
Compare systems on the same cases
Pair results by case. Count cases where A passes and B fails, and vice versa. Use a paired analysis appropriate to the outcome, such as McNemar’s test for paired binary outcomes under its assumptions. For continuous case-level scores, a paired bootstrap can estimate the uncertainty in the difference if the sampling unit is correct.
Match prompts, tool access, reference versions, budgets, and environment unless the intended comparison deliberately changes them. Report cost and latency as separate dimensions. Do not hide failures with post hoc exclusions.
Repeated attempts
A first-attempt success rate asks whether the initial run succeeds. Best-of-k success asks whether at least one of k attempts succeeds. All-k reliability asks whether every attempt succeeds. These answer different operational questions. Best-of-k results can overstate a workflow that users only run once.
The conventional sampled pass@k estimator assumes n candidate solutions with c correct: 1 - C(n-c, k) / C(n, k) for valid k. Do not calculate it from one run per case or imply independent retries when the attempts share state.
Exercise: model A improves average correctness but doubles critical omissions. Write the decision report without hiding the critical failures in the mean. Specify which further evidence is needed.