38. Build and inspect an agent from the ground up

Published

October 3, 2026

An agent system combines a model with a runtime that supplies context, routes proposed actions, records observations, and decides when work stops. The model proposes; application code executes. Understanding that separation makes a coding assistant, a research assistant, and a surveillance assistant easier to inspect even when their interfaces look different.

NoteLearning objectives
  • Trace a request through context, proposals, authorization, and observations.
  • Run a complete bounded loop and reproduce incomplete execution.
  • Identify what a provider adapter changes and what the runtime must enforce.
TipTL;DR

An agent proposes actions; the runtime validates and executes them. Run the bounded loop, inspect its observations, and reproduce denied or budget-exhausted outcomes. A final message is not an execution receipt or proof of correctness.

Introduction

Start from the tool authorization in lesson 34. Add a repeated decision loop and a terminal result. The program below is fully runnable using authored policies; replacing that policy with a language-model adapter creates a new system whose behavior must be evaluated separately.

The request-to-result architecture

Human task and authenticated scope
        |
        v
Context builder: task + eligible sources + observations
        |
        v
Model adapter: structured action proposal or final answer
        |
        v
Harness: validate proposal, enforce budget, record event
        |
        +-- final --> independent outcome and evidence checks
        |
        v
Trusted authorization --> permitted tool --> actual observation
        |                                      |
        +---------------- next turn -----------+

The context builder should preserve source identity and time. The adapter handles provider requests and structured responses. The harness controls execution. The tool service checks its own authorization. The evaluator grades artifacts and state rather than relying on the agent’s self-assessment.

A fixed workflow follows predetermined steps. A model-directed agent chooses its next actions within the available interface. Anthropic’s Building effective agents explains that architectural distinction. Its original tooling discussion is historical; the later Managed Agents engineering account separates session logs, harnesses, and execution environments. That separation helps recover from failure without making credentials accessible to generated code.

Run a complete small loop

python3 foundations/agent_loop.py

Expected behavior: the policy first requests read_completeness for synthetic-a; the runtime returns “missing source date”; the next turn produces “Manual review: missing source date.” The saved result has status complete, one attempted tool, and one observation.

from __future__ import annotations

import json
from dataclasses import asdict, dataclass
from typing import Protocol

from tool_runtime import Call, execute


@dataclass(frozen=True)
class Observation:
    tool: str
    record_id: str
    text: str


@dataclass(frozen=True)
class Final:
    text: str


class Policy(Protocol):
    def propose(self, observations: tuple[Observation, ...]) -> Call | Final: ...


@dataclass(frozen=True)
class Run:
    status: str
    answer: str
    observations: tuple[Observation, ...]
    attempted_tools: tuple[str, ...]
    error: str | None


def run(policy: Policy, allowed_ids: set[str], max_turns: int = 3) -> Run:
    if type(max_turns) is not int or max_turns < 1:
        raise ValueError("max_turns must be a positive integer")
    observations: tuple[Observation, ...] = ()
    attempted: tuple[str, ...] = ()
    for _ in range(max_turns):
        proposal = policy.propose(observations)
        if isinstance(proposal, Final):
            if not proposal.text.strip():
                raise ValueError("empty final answer")
            return Run("complete", proposal.text, observations, attempted, None)
        attempted += (proposal.name,)
        try:
            result = execute([proposal], allowed_ids)[0]
        except PermissionError as error:
            return Run("denied", "", observations, attempted, str(error))
        observations += (Observation(result["tool"], result["record_id"], result["observation"]),)
    return Run("budget_exhausted", "", observations, attempted, "no final answer within turn budget")


class InspectPolicy:
    def propose(self, observations: tuple[Observation, ...]) -> Call | Final:
        if not observations:
            return Call("read_completeness", "synthetic-a")
        return Final("Manual review: " + observations[-1].text)


class ScriptedPolicy:
    # Authored proposals exercise the harness. This is not a language model.
    def __init__(self, proposals: tuple[Call | Final, ...]) -> None:
        if not proposals:
            raise ValueError("empty script")
        self.proposals = proposals
        self.index = 0

    def propose(self, observations: tuple[Observation, ...]) -> Call | Final:
        if self.index >= len(self.proposals):
            raise ValueError("script ended without a final answer")
        proposal = self.proposals[self.index]
        self.index += 1
        return proposal


def main() -> None:
    print(json.dumps(asdict(run(InspectPolicy(), {"synthetic-a"})), sort_keys=True))


if __name__ == "__main__":
    main()

Policy is a typed interface. InspectPolicy and ScriptedPolicy are complete deterministic implementations, not model substitutes presented as model performance. Call and execute come from the earlier runtime module. Run retains both the answer and the execution evidence.

Trace each boundary in the code

The policy receives observations, not a mutable permission object. allowed_ids comes from the caller. Each proposed tool passes through execute; the policy cannot grant itself an additional tool by naming it.

A permission denial becomes a terminal denied result with its reason. Unknown records and implementation errors raise rather than being quietly turned into successful answers. A final answer has to contain text. When the turn budget ends before a final answer, the result is budget_exhausted, with an empty answer and an explicit error.

The turn budget includes the final-answer turn. With a budget of one, a policy that first reads a record cannot subsequently finish. That is intentional and tested. Production resource controls may separately track model turns, tool calls, elapsed time, and billed tokens.

An agent can return a final answer that is unsupported or incorrect. Completion means the protocol reached its terminal state, not that the task was done well. The security scorer evaluates usefulness independently. Domain graders must do the same for clinical or scientific claims.

Replace the policy with a provider adapter

The implementation checklist is concrete:

  1. Serialize the task, observations, and model-visible sources into the provider’s documented message format. Preserve IDs; do not expose evaluator references.
  2. Supply the approved tool schema. Request an action with exact tool name and arguments, or a final answer.
  3. Save the raw response with model identifier, request ID when supplied, tool schemas, and inference settings.
  4. Parse into Call or Final. Reject unknown tools, unexpected fields, invalid argument types, ambiguous multiple actions, and empty final text.
  5. Route the proposal through the existing authorization layer. Do not execute text as shell commands merely because it looks like code.
  6. Preserve explicit timeout, provider error, cancellation, and parsing-failure statuses. Bound retries; reconcile mutations with uncertain outcomes.
  7. Compare the new adapter on the same task contract, plus fresh cases and tool failures.

A provider response that contains a tool call has not executed the tool. Execution belongs to the application or the managed runtime. Read lesson 29 for API and concurrency boundaries, and lesson 21 for durable runner state.

No credentials are needed for this offline exercise. Adding a live endpoint requires an approved account, current interface documentation, and explicit cost limits. The download does not contain an unfinished API connector.

Understand Codex, Claude Code, and Cursor by inspecting the same surfaces

Surface What to inspect in each product Evidence of understanding
Task context Repository instructions, selected files, retrieved material Explain what the agent could see
Model interaction Selected model and available settings Record a configuration rather than a product name alone
Execution Local process, remote environment, or managed sandbox Identify where edits and commands actually run
Authority Approved tools, paths, network destinations, identity Reproduce an allowed action and a denied action
State Session transcript, files, memory, checkpoints Recover a failure without losing provenance
Completion Diff, tests, artifact, destination receipt Verify an outcome beyond the final message

Use lesson 4 for the canonical coding-agent prompt workflow. For current controls, use Codex’s safety account, Claude Code security, and Cursor agent security. Do not assume that a control described for a hosted agent also applies to its desktop or local terminal counterpart.

Personal agents: Grok Bot, Muse, and Instinct

The following are verified public product descriptions, checked October 3, 2026. They identify useful architectures to study; they are not independent security audits or reconstructions of proprietary source code.

Product and primary source Documented design Practical evaluation question
Grok Bot overview Persistent cloud computer with browser, files, terminal, and connectors; an account’s Bots share that computer Can another Bot in the same account see task files or inherit a logged-in session? Is that sharing appropriate for the task?
Meta Muse safety engineering Runtime cell separated from credential-capable services; Sentinel governs connector actions and network egress Does a harmless unauthorized proposal get denied outside the main agent, and does legitimate work still complete?
Meta Muse design Public account of personal-agent design and interactions Can the user inspect what happened, correct a mistaken assumption, and bound the next action?
Instinct official site Personal assistant accessed through text or calls, connected to applications and devices What scope is granted, where does an action execute, and what receipt proves it completed?

Grok Bot’s documentation explicitly describes sharing across Bots within one account. Treat that as part of the access model when designing an eval. It differs from isolating each worker in a separate fresh environment. A task that requires hidden references must keep them outside the shared computer.

Meta’s account gives a concrete example of the proposal-versus-authority distinction: the main agent proposes an action and a separate permission component authorizes it. The engineering account describes the launch design; confirm the current implementation and configuration before treating any control as present in a tested run. An architectural description is not a measured prompt-injection resistance score.

Instinct’s homepage identifies interaction and integration surfaces but does not disclose enough detail to reconstruct its complete execution harness. Apply the inspection table using official documentation or authorized source access. Do not infer its internal model, training process, or isolation mechanism from a demonstration video.

The name Docty also appears in healthcare products, so the intended agent requires an exact official URL before product-specific instructions can be reliable. The generic architecture and inspection exercise remain usable meanwhile. No undocumented internals are assumed for any product.

Context, memory, and delegation

The context window is what the model sees for the current request. A durable session log can retain more history than fits in that window. A summary is a lossy transformation of earlier evidence; keep the original events available when auditability matters.

Persistent memory needs provenance and scope. A saved preference cannot silently become a permission, and a prior patient’s fact cannot become another patient’s context. Test revocation and identity changes. Retrieval should select eligible evidence; memory should not override a current source without a declared policy.

Multiple agents add scheduling, ownership, and reconciliation problems. Begin with the small loop and add delegation only when independent tasks justify it. Evaluate duplicate edits, conflicting conclusions, hidden-reference exposure, and costs across the whole team. Several agents agreeing is not independent evidence if they share the same flawed source or judge.

Ground-up build exercise

Use a coding agent to extend the kit with a second read-only synthetic tool. Provide the allowed record IDs, tool schema, expected observation, and denial cases. Require tests for unknown IDs, budget exhaustion, unsupported final claims, and recovery after an explicit failure. Inspect the diff before accepting it.

The stopping artifact is a diagram whose arrows match executed code, a saved trace, a supported answer, and reproduced failure paths. That is enough to explain a small agent system from the ground up. Production deployment adds durable state, operational monitoring, access management, and domain validation.

Checkpoint: change the turn budget to one in a local copy, reproduce the incomplete result, and explain why the final answer cannot be treated as the tool receipt.