The minimum engineering you need to understand
TL;DR
Understand functions, types, state, side effects, errors, and tests well enough to inspect generated code. Focus on Python and the invariants a change must preserve.
You do not need to memorize a programming language before running the first exercise. You do need enough engineering literacy to recognize what a coding agent changed and why it might fail. Start with a single stack, Python for this course, rather than sampling several languages superficially.
The system beneath an answer
A model receives tokens and emits tokens. An application adds prompts, retrieval, memory, tool definitions, permissions, and user-interface behavior. An agent adds an action loop: observe, choose an action, call a tool, receive a result, continue, stop. A harness orchestrates those steps and records them. The system under test is that configured combination, not the model name alone.
| Concept | What you must be able to inspect |
|---|---|
| Function | Inputs, returned value, side effects, and failure behavior |
| Type | Whether a field can be missing, null, numeric, textual, or categorical |
| Dependency | What package is required and which version is installed |
| Process | What runs, its exit status, environment, and resource limits |
| API | Request schema, response schema, authentication, errors, and rate limits |
| Database | How records are identified, updated, queried, and isolated |
| Test | Which behavior is asserted and which alternative behaviors remain untested |
| Build | How source becomes the runnable or publishable artifact |
An exit code of zero usually signals successful process completion. It does not prove that the program performed the intended task. A test command can exit successfully after collecting no relevant tests if its configuration is wrong. Read the output and verify that expected tests were discovered.
Read this function
def proportion(numerator: int, denominator: int) -> float:
if denominator <= 0:
raise ValueError("denominator must be positive")
if numerator < 0 or numerator > denominator:
raise ValueError("numerator outside valid range")
return numerator / denominator
The annotations express an interface. The checks enforce a narrower mathematical contract. Ask an agent why proportion(3, 2) should fail and why returning zero for proportion(0, 0) would conceal a missing denominator. Then ask it to write tests independently from the function body, based on the contract.
Python's bool is a subtype of int; annotations alone do not reject a Boolean count. If the function is a public data boundary, explicit runtime validation may be necessary. Distinguish typed internal code from untrusted external inputs.
State and side effects
A pure computation transforms inputs without modifying them. A side effect changes a file, sends a request, updates a database, or mutates a shared object. Side effects create evaluation risk: repeated runs may see different starting conditions. A clinical message draft and a sent message are different outcomes. A database read and a committed write are different permissions.
For each tool, list the side effects. Reset state between trials. Use synthetic test accounts. For an evaluator, a dry-run switch is useful only if you have verified that it suppresses every relevant side effect.
Concepts that expose weak engineering
Recognize invariants (conditions that must remain true), idempotency (repeating an operation does not add unintended effects), race conditions (timing changes outcomes), transactions (related updates succeed or fail together), and backward compatibility (existing users retain valid behavior). You need to reason about these before judging architecture, even if an agent writes the code.
Exercise: an agent fixes duplicate alerts by removing all duplicates from a database every morning. Explain why that may leave a window for duplicate messages and why deduplicating the send operation with a stable event key could be stronger. Identify when an update to an existing alert should legitimately create a new event.