25. Glossary and primary-source reading map
Use this glossary to orient yourself while working. Learn terms through the artifact they describe, not by memorizing definitions in isolation.
- Use core terminology accurately.
- Find primary documentation for changing interfaces.
- Distinguish sources from teaching fixtures.
Use the glossary while building artifacts and consult primary documentation for changing interfaces. Sources guide learning; teaching exercises are not published research results.
| Term | Operational meaning |
|---|---|
| Benchmark | A declared set of tasks and scoring procedures used for comparison |
| Eval | A measurement procedure for a specified system behavior or claim |
| Harness | The execution and recording machinery around the tested system |
| Scaffold | Tools, prompts, control flow, and other support around a model |
| Reference standard | The policy or evidence used to adjudicate correctness |
| Grader | A procedure that maps observed artifacts to scores or judgments |
| Oracle | A source of expected behavior, which may itself be fallible |
| Trace | The observable sequence of messages, actions, and intermediate results |
| Holdout | Cases withheld from tuning, until inspected or used for development |
| Leakage | Information entering evaluation that invalidates the intended test |
| Contamination | Prior exposure or overlap that undermines a capability inference |
| Calibration | Aligning judgments or predicted confidence with reference observations |
| Ablation | Changing or removing one component to study its contribution |
| Regression | A previously acceptable behavior that becomes worse after a change |
| Sandbox | A controlled execution environment with declared isolation properties |
| RAG | Retrieval-augmented generation, adding retrieved material to model input |
| Groundedness | Whether output claims are supported by the supplied evidence |
| Robustness | Behavior under relevant perturbations or stress conditions |
| Coverage | Which declared cases, conditions, or outcomes were examined |
Reading map
These sources were located during preparation on October 2, 2026. Documentation and live services can change. Recheck exact interfaces before implementing an integration. Course exercises and proposed procedures are teaching designs, not published experimental results.
- Inspect official documentation: consult for framework concepts and execution interfaces.
- Anthropic: Demystifying evals for AI agents: consult for the short conceptual orientation cited in chapter 11.
- OpenAI: Separating signal from noise in coding evaluations: consult for benchmark-validity scrutiny cited in chapter 13.
- Codex learning resources: current official product-learning entry point.
- Claude Code engineering guidance: official engineering context; verify current interfaces separately.
- Cursor documentation: current official editor-agent documentation.
- TRIPOD-LLM, Nature Medicine: reporting guideline. Search metadata was available; direct full text was not accessible during preparation.
- NIST Generative AI Profile: governance and risk-management reference.
- HHS de-identification guidance: primary U.S. privacy guidance for the limited point cited in chapter 20.
- Evaluating epidemic forecasts in an interval format, preprint: primary methodological reference for the interval-scoring discussion.
What to do next
Return to chapter 8 and run the kit. Write your first report. Then complete the coding lab and one domain task contract. The course becomes useful when those artifacts reveal what you understand, what you can verify, and what still requires expert review.