32. Mathematics and a complete small training experiment
A training loop converts a chosen error measure into parameter updates. A simple classifier makes that process visible without a GPU or a deep-learning framework. The invented task below predicts whether a number belongs to the positive or negative half of a synthetic distribution; it has no clinical interpretation.
- Explain logits, loss, gradients, and parameter updates.
- Keep test cases separate from training.
- Run a controlled label-shift experiment.
Train a small classifier and inspect the parameters, losses, and withheld cases. Reducing training loss does not establish generalization or clinical usefulness.
Introduction
A scalar is one number. A vector is an ordered collection of numbers. A matrix arranges numbers in rows and columns. A tensor generalizes this to more dimensions. For a batch of B cases with D features, an input array commonly has shape B by D. A linear classifier combines the features using weights of length D and a bias. The output must still align with the B case identities.
For one feature, the logit is z = weight * x + bias. The sigmoid maps that score to a number between zero and one. The label is zero or one. Binary cross-entropy penalizes an incorrect confident prediction more strongly than an uncertain prediction. A reported probability is not automatically calibrated to a new population.
The derivative of this loss with respect to the logit is predicted_probability - label. For one example, the weight gradient is that residual multiplied by x; the bias gradient is the residual. Averaging these gradients over the training batch gives the update used below. Gradient descent subtracts the gradient multiplied by a learning rate.
Run the complete experiment
From the extracted starter-kit folder:
python3 foundations/train_model.py
python3 -m unittest discover -s tests -vfrom __future__ import annotations
import json
import math
from dataclasses import dataclass
@dataclass(frozen=True)
class Row:
case_id: str
x: float
y: int
def validate(rows: list[Row]) -> None:
if not rows:
raise ValueError("empty dataset")
if len({row.case_id for row in rows}) != len(rows):
raise ValueError("duplicate case ID")
for row in rows:
if not row.case_id or not math.isfinite(row.x) or row.y not in (0, 1) or isinstance(row.y, bool):
raise ValueError("invalid row")
def sigmoid(z: float) -> float:
if not math.isfinite(z):
raise ValueError("nonfinite logit")
if z >= 0:
return 1.0 / (1.0 + math.exp(-z))
exp_z = math.exp(z)
return exp_z / (1.0 + exp_z)
def loss(rows: list[Row], weight: float, bias: float) -> float:
validate(rows)
values = [max(weight * r.x + bias, 0.0) - r.y * (weight * r.x + bias)
+ math.log1p(math.exp(-abs(weight * r.x + bias))) for r in rows]
if not all(math.isfinite(value) for value in values):
raise ValueError("nonfinite loss")
return sum(values) / len(rows)
def train(rows: list[Row], steps: int = 200, rate: float = 0.1) -> tuple[float, float]:
validate(rows)
if isinstance(steps, bool) or steps <= 0 or not math.isfinite(rate) or rate <= 0:
raise ValueError("invalid optimization settings")
weight = bias = 0.0
for _ in range(steps):
residuals = [sigmoid(weight * r.x + bias) - r.y for r in rows]
weight -= rate * sum(error * row.x for error, row in zip(residuals, rows, strict=True)) / len(rows)
bias -= rate * sum(residuals) / len(rows)
return weight, bias
def experiment(train_rows: list[Row], test_rows: list[Row]) -> dict[str, float | int]:
validate(train_rows)
validate(test_rows)
if {r.case_id for r in train_rows} & {r.case_id for r in test_rows}:
raise ValueError("train/test ID overlap")
weight, bias = train(train_rows)
correct = sum(int(sigmoid(weight * r.x + bias) >= 0.5) == r.y for r in test_rows)
return {"train_before": loss(train_rows, 0.0, 0.0), "train_after": loss(train_rows, weight, bias),
"test_loss": loss(test_rows, weight, bias), "test_correct": correct,
"test_total": len(test_rows), "weight": weight, "bias": bias}
def main() -> None:
# Synthetic sign classification, deliberately simple. No patient predictions.
training = [Row("tr1", -2.0, 0), Row("tr2", -1.0, 0), Row("tr3", 1.0, 1), Row("tr4", 2.0, 1)]
testing = [Row("te1", -0.5, 0), Row("te2", 0.5, 1)]
print(json.dumps(experiment(training, testing), indent=2, sort_keys=True))
if __name__ == "__main__":
main()The experiment prints the initial and final training loss, a held-out test loss, the learned weight and bias, and the exact test numerator and denominator. Inspect the actual output instead of memorizing a claimed benchmark score. The sample was constructed to be easily separable, so success says little about difficult data.
Understand each boundary
validate rejects empty data, duplicate IDs, invalid labels, and nonfinite features. The train/test overlap check is necessary for this exercise but does not detect every form of leakage. Two different IDs could refer to the same patient or near-duplicate document. Real splits need the relevant grouping unit and, for forecasting, the correct time boundary.
The numerically stable sigmoid handles positive and negative logits separately. The loss uses a stable expression based on log1p, avoiding a direct logarithm of a probability that rounded to zero. Numerical stability is an engineering property; it does not make the dataset or objective scientifically valid.
train starts both parameters at zero. Each iteration computes residuals using only training rows, averages the gradients, and updates the parameters. experiment evaluates the resulting parameters on test rows. It does not use test performance to select the learning rate or training duration.
The default step count and learning rate are instructional settings, not recommended clinical settings. A validation split would be needed to choose settings empirically. A final acceptance set should remain untouched until the selection process is complete.
Compare against a baseline and challenge generalization
For the supplied balanced two-case test set, an always-zero rule gets one label right. This follows directly from the authored labels. Compare the trained model against that rule, reporting the two-case denominator. Such a tiny test cannot support a reliable population-level performance estimate.
Now invert the test labels while preserving the features and IDs. Run again. Training behavior remains the same because test labels do not update parameters, but test performance changes. This is a controlled illustration of a changed relationship between inputs and targets, not evidence about actual clinical distribution shift.
Next introduce an overlapping ID. The experiment must reject the run. Introduce float("nan") as a feature through a Python test. It must reject the row before reporting loss. These negative controls verify that the experiment does not convert a defective input into a measured result.
What changes in a neural network
A neural network replaces the single linear rule with compositions of parameterized transformations and nonlinear activations. Backpropagation applies the chain rule to compute gradients through those transformations. An optimizer uses those gradients to update parameters. Larger models increase computational and data demands; they do not remove the need for sound splits and task-specific evaluation.
In PyTorch, a typical loop clears accumulated gradients, computes a forward pass and loss, calls backpropagation, and steps the optimizer. Evaluation should use the appropriate evaluation mode and disable unnecessary gradient tracking. Train/eval mode and gradient tracking are different controls. The official PyTorch foundations sequence connects tensors, datasets, autograd, optimization, and saving a model.
Checkpoint: explain the residual, gradient, learning rate, and separation of training from test data. Identify what changes in the inverted-test-label experiment and what remains fixed. A training loss decrease alone is not a deployment recommendation.