31. Python and data contracts from a runnable example
Python literacy makes generated code inspectable. The practical starting point is a small program whose inputs, transformations, outputs, and failure conditions can all be explained. The example below validates invented evaluation records before counting their labels.
- Trace data from JSON through a typed record to an output.
- Reject duplicate IDs and invalid schema fields.
- Run and explain failure-path tests.
Read a complete typed Python validator, run it, and reproduce rejected inputs. Data validation precedes counting; annotations alone do not validate arbitrary JSON.
Introduction
A variable names a value. A type describes the values permitted at a boundary. A function accepts inputs, performs work, and returns a result or raises an error. A module is a file of related definitions. Importing a module makes its definitions available; its if __name__ == "__main__" block runs only when that file is executed directly.
| Python object | Evaluation use | Important distinction |
|---|---|---|
str |
Case ID, source text, label | "2" is text, not a count |
int |
Source day, attempt count | Boolean values need explicit exclusion at JSON boundaries |
float |
Loss, probability, elapsed time | Reject nonfinite values where meaningful numbers are required |
list |
Ordered cases or responses | Order and duplicate entries are preserved |
dict |
Fields keyed by name | A missing key differs from a present key with a null value |
set |
Unique IDs and membership checks | Conversion to a set discards duplicates |
| Dataclass | A typed record | Type annotations alone do not validate arbitrary input |
An R vector and a Python list are not interchangeable numerical abstractions. A Python list can hold different object types; elementwise numerical operations normally use arrays or explicit iteration. Python uses zero-based indexing: the first item is items[0]. The slice items[1:3] includes indices 1 and 2, excluding 3. These details matter when a batch unexpectedly loses its first or last case.
Run and inspect the program
Download the starter kit from the Practice Lab and extract it. From its eval-starter-kit directory, run:
python3 --version
python3 foundations/data_contract.py
python3 -m unittest discover -s tests -vThe program uses only the standard library. No API key, package installation, patient data, or model access is required. The authored example contains one complete record and one review record. Its JSON output reports those counts and the validated records.
from __future__ import annotations
import json
from collections import Counter
from dataclasses import asdict, dataclass
@dataclass(frozen=True)
class Case:
case_id: str
label: str
source_day: int
def parse_cases(raw: object) -> list[Case]:
if not isinstance(raw, list) or not raw:
raise ValueError("expected a nonempty list")
cases: list[Case] = []
seen: set[str] = set()
for item in raw:
if not isinstance(item, dict) or set(item) != {"id", "label", "source_day"}:
raise ValueError("incorrect case schema")
case_id: object = item["id"]
label: object = item["label"]
day: object = item["source_day"]
if not isinstance(case_id, str) or not case_id.strip() or case_id != case_id.strip():
raise ValueError("invalid case ID")
if case_id in seen:
raise ValueError("duplicate case ID")
if not isinstance(label, str) or label not in {"complete", "review"}:
raise ValueError("unknown label")
if isinstance(day, bool) or not isinstance(day, int) or day < 0:
raise ValueError("invalid source day")
seen.add(case_id)
cases.append(Case(case_id, label, day))
return cases
def main() -> None:
# Invented records describe data completeness, not clinical severity.
raw: object = json.loads('[{"id":"a","label":"complete","source_day":2},'
'{"id":"b","label":"review","source_day":3}]')
cases = parse_cases(raw)
print(json.dumps({"cases": [asdict(case) for case in cases],
"counts": dict(Counter(case.label for case in cases))}, sort_keys=True))
if __name__ == "__main__":
main()Read it as an execution trace
json.loads converts JSON text into Python objects. The object annotation deliberately means the parser has not yet earned a narrower type. The program first requires a nonempty list, then validates every item. Exact key matching rejects misspellings and unexpected fields rather than quietly ignoring them.
Each ID must be a nonblank string without surrounding whitespace. seen tracks IDs already encountered, so duplication becomes an error before a counter can double-count a case. Labels come from a declared vocabulary. A source day must be a nonnegative integer. The Boolean exclusion matters because Python treats bool as a subclass of int.
Only validated values enter a Case. frozen=True prevents ordinary assignment to its fields after construction. It does not make arbitrary nested objects immutable; these fields are simple scalar values. Counter counts the labels. asdict converts the dataclass records back into dictionaries for JSON export.
Make the program fail deliberately
Change the second ID to a, keeping both records. The program must raise a duplicate-ID error and exit unsuccessfully. Then restore it and change source_day to true in the JSON. That must also fail. Finally, add an unexpected field. A failure here is correct behavior, not an inconvenience to suppress.
For a real dataset, preserve the raw source, the validation report, and the versioned accepted data separately. Do not turn malformed records into empty dictionaries or drop them while reporting a complete cohort. A pipeline may support quarantined records, but that policy must be explicit and its denominator visible.
Ask a coding agent for one bounded extension
Add an optional source_version field through a separate explicit schema revision.
Keep the current parser strict for its existing schema.
Add tests for accepted versions, unknown versions, duplicate IDs, and missing fields.
Do not add a dependency or silently default an unknown version.
Run the tests and explain how callers select the schema.
The review question is whether the new boundary preserves the old contract and makes the revision visible. A persuasive explanation without an executed failure test is insufficient.
Checkpoint: explain why duplicate detection happens before aggregation, why annotations do not replace parsing, and why missing data cannot be silently converted into success. Then run a valid input and reproduce two rejected inputs.