33. Retrieval and evidence grounding you can test

Published

October 3, 2026

Retrieval supplies candidate evidence to an answering system. It does not establish that the selected evidence is current, applicable, or correctly interpreted. A transparent lexical baseline makes retrieval, time filtering, and literal support separate testable operations.

NoteLearning objectives
  • Filter sources by the relevant decision time.
  • Measure retrieval independently from answer support.
  • Reproduce wrong-source and no-match cases.
TipTL;DR

Evaluate time eligibility, retrieval relevance, and answer support separately. A literal quotation check establishes a narrower property than clinical correctness.

Introduction

A retrieval-augmented system typically ingests sources, divides them into usable units, indexes those units, retrieves candidates for a query, and provides selected context to a model. The answering model then generates an output that requires its own evaluation. A source can be retrieved successfully and still be misquoted or misapplied.

The local exercise uses authored surveillance-policy sentences. These are invented policies for testing information flow, not public-health recommendations. The later version must be unavailable at an earlier decision time, even if it matches the query very well.

Run the baseline

python3 foundations/retrieval.py
python3 -m unittest discover -s tests -v
from __future__ import annotations

import json
import re
from dataclasses import dataclass


@dataclass(frozen=True)
class Document:
    doc_id: str
    published_day: int
    text: str


def terms(text: str) -> set[str]:
    return set(re.findall(r"[a-z0-9]+", text.lower()))


def retrieve(query: str, docs: list[Document], as_of: int, limit: int = 2) -> list[Document]:
    if not terms(query) or as_of < 0 or limit < 1:
        raise ValueError("invalid retrieval request")
    if len({d.doc_id for d in docs}) != len(docs):
        raise ValueError("duplicate document ID")
    if any(not d.doc_id or not d.text.strip() or d.published_day < 0 for d in docs):
        raise ValueError("invalid document")
    # Lexical overlap is an inspectable baseline, not a semantic embedding.
    scored = [(len(terms(query) & terms(d.text)), d) for d in docs if d.published_day <= as_of]
    ranked = sorted(scored, key=lambda pair: (-pair[0], pair[1].doc_id))
    return [d for score, d in ranked if score > 0][:limit]


def quoted_answer(doc_id: str, quote: str, context: list[Document]) -> bool:
    # Verifies literal quotation and source identity, not medical correctness.
    return bool(quote.strip()) and any(d.doc_id == doc_id and quote in d.text for d in context)


def main() -> None:
    docs = [Document("policy-v1", 2, "Missing surveillance counts require review."),
            Document("policy-v2", 12, "Missing surveillance counts can be imputed."),
            Document("other", 1, "Archive documents by source ID.")]
    context = retrieve("missing surveillance counts", docs, as_of=5)
    print(json.dumps({"retrieved": [d.doc_id for d in context],
                      "supported": quoted_answer("policy-v1", "require review", context),
                      "future_supported": quoted_answer("policy-v2", "can be imputed", context)}, sort_keys=True))


if __name__ == "__main__":
    main()

The query has overlapping terms with both policy versions. At day 5, only policy-v1 is eligible. The program verifies a supplied exact quotation from that context and rejects a quotation attributed to the future version. It never calls a language model.

What the baseline computes

terms lowercases and extracts alphanumeric tokens. It returns a set, so repeated words do not increase this overlap score. Each eligible document gets a score equal to the number of shared query terms. Positive scores are sorted in descending order, with document ID providing a deterministic tie-break. The limit is applied after filtering and ranking.

This is deliberately simpler than TF-IDF, BM25, or embedding retrieval. It can fail when synonyms share no tokens, when an irrelevant document repeats relevant terms, or when negation changes meaning. Its virtue is inspectability, not superior performance.

An embedding represents text numerically. A similarity metric ranks representations. A reranker can reconsider candidate relevance. None of these operations proves source authority. A current guideline for a different population can still be unsuitable evidence for a clinical answer.

Design two evaluations

For retrieval, define the eligible source universe, query, acceptable documents, and decision time. Report whether the required evidence appears among the first k retrieved items. If multiple documents are acceptable, retain that reference set. A query with no eligible relevant document is an important abstention case, not an invitation to invent an answer.

For the answer, assess source identity, claim support, omitted qualifications, population applicability, and handling of conflicting evidence. The exercise’s quoted_answer verifies only that a nonblank string occurs literally in the attributed source. It does not establish entailment for a broader paraphrase, completeness, medical correctness, or appropriateness for a patient.

Stress the components independently

  1. Query using a synonym absent from the source text. The lexical retriever may return nothing. Record this as a retrieval limitation.
  2. Change as_of from 5 to 15. Both versions become eligible. Observe that overlap ranking alone does not resolve version precedence.
  3. Attribute the correct quotation to the wrong document ID. Literal source validation must fail.
  4. Supply a blank quotation. Validation must fail.
  5. Add two documents with the same ID. Retrieval must reject the corpus rather than choose one silently.

To handle superseded policies, add explicit effective dates and supersession relationships instead of assuming the latest publication is always controlling. Publication time, availability time, and effective time can differ. The evaluation must state which of them determines eligibility.

Extend with a model only after the apparatus works

Provide the retrieved source IDs and text to a model under a bounded task: quote the applicable completeness rule or state that eligible support is missing. Preserve the exact prompt, context, response, source versions, and runtime configuration. Grade against the reference without exposing it to the tested model.

Compare the same cases with and without retrieval. A useful experiment reports retrieval failures separately from failures after correct evidence was supplied. Otherwise a low final score cannot identify which intervention is appropriate.

Checkpoint: demonstrate a future-source exclusion, wrong-source rejection, and no-match result. Explain why a perfect quotation check does not establish a clinically correct answer.