11. Evaluate agents, tools, retrieval, and memory

Published

October 3, 2026

An agent’s outcome depends on what it can see and do. Measure the configured workflow, including tools, retrieval, permissions, and persistent state. If you change those, you have changed the tested system.

NoteLearning objectives
  • Specify tools, retrieval, memory, and permissions.
  • Separate final-answer checks from trajectory checks.
  • Test whether evidence supports the output.
TipTL;DR

Evaluate the configured model, tools, retrieval, memory, and permissions together. Separate final outcomes from trajectories and inspect whether retrieved evidence truly supports claims.

Outcome and trajectory

An outcome is the final state: a correct patch, a faithful summary, or a saved draft. A trajectory is the observable sequence of messages and tool actions. A correct outcome can follow an unacceptable trajectory, such as obtaining an answer from a forbidden reference. A safe trajectory can end in an incomplete outcome.

Record both. Avoid treating private chain-of-thought as required evidence. Tool calls, visible rationales, intermediate artifacts, and verified state changes provide actionable evaluation material.

Retrieval evaluations

Separate whether the relevant source was retrieved from whether the final answer used it correctly. Retrieval recall requires a defined relevant-document set and denominator. Faithfulness requires checking whether the answer’s claims are supported by the supplied sources. Citation presence alone proves neither.

A useful test deliberately supplies an older document and a current correction. Does the system recognize which version governs the answer? Another supplies a document containing irrelevant instructions. Does the system treat it as evidence rather than authority over its behavior?

Tool behavior

Check argument validity, tool-selection appropriateness, permission checks, error handling, and postcondition verification. A request to create a draft should not send it. A read-only review should not alter the chart. If a tool returns “success,” inspect whether the intended state actually exists.

Represent tools with synthetic or mock services when possible. Make errors explicit: access denied, no results, stale results, malformed records, timeout, and partial completion. A mock should match the relevant interface and failure semantics; it is not evidence that a real integration works.

Memory and long conversations

Use paired cases with and without relevant prior context. Test whether the agent remembers an allergy, preserves a change in permission, and stops using a superseded instruction. Test conflicting context explicitly. A long transcript can induce omission even when a short version passes.

Reset memory between unrelated cases. Record whether memory persists across trials. Otherwise, later cases may benefit from earlier answers or leak information between users.

Framework orientation

Inspect documentation describes tasks composed of datasets, solvers, and scorers, with logs and sandbox support. Use its official guide when moving beyond the local teaching scorer. The conceptual mapping is: cases become a dataset, your system interaction becomes a solver, and your rubric becomes a scorer. Framework adoption does not validate your reference standard.

Anthropic’s agent-evaluation overview provides a short orientation to outcomes, transcripts, and grader categories. The procedures in this course are teaching recommendations to apply to your own task, not a reproduction of that article.

Exercise: design a tool-using trial matcher with one denied tool call. Specify the expected behavior after denial and how you would detect that it fabricated access or eligibility evidence.