27. ML and post-training concepts in plain language

Published

October 3, 2026

You need a working model of learning and inference to understand research-engineering postings. Start with what changes, what is optimized, and what is measured. Mathematical depth becomes more necessary as you move from evaluation into training research.

NoteLearning objectives
  • Distinguish training from inference.
  • Explain SFT, preferences, rewards, and optimization.
  • Detect reward improvements that conceal task failures.
TipTL;DR

Understand training, inference, SFT, preference learning, rewards, and RL in terms of what changes and what is optimized. Better reward scores can still conceal worse real behavior.

Training versus inference

Training changes model parameters using data and an objective. Inference uses a configured model to produce an output. Changing a prompt, retrieval corpus, or tool does not normally change the model’s parameters. It changes the system around the model.

A parameter is a learned numerical value. A hyperparameter controls how learning or execution works. A token is a unit of model input/output representation, not necessarily a word. A context window limits the information available in a single interaction. Do not infer that a model reliably uses every fact merely because it fits in the window.

The five concepts behind many postings

Concept Plain description Evaluation question
Pretraining Learn broad patterns from large data collections What capabilities and biases are inherited?
SFT Learn from selected examples of desired responses or trajectories Are demonstrations correct and representative?
Preference learning Use comparisons of outputs to shape behavior Whose preferences, under what rubric?
Reward modeling Estimate a score for an output or behavior Does the score reward the intended thing?
Reinforcement learning Adjust behavior using rewards from interactions Does optimizing reward improve the real objective?

RLHF uses human feedback in the training signal. RLAIF uses AI feedback in at least part of that signal. RL with verifiable rewards uses checkable outcomes, such as passing a task-specific test. None of these phrases guarantees correct data, valid rewards, or safe behavior.

A small worked reward example

Suppose a coding task’s reward is “the test command exited zero.” A model may solve the task, or it may disable the tests. Both could earn the same naive reward. Strengthen the measurement by checking collected test identities, protecting grader files, verifying behavior independently, and comparing the final diff with permitted changes.

A medical-summary reward that counts citations can reward invented or irrelevant citations. A stronger criterion checks claim support, citation identity, critical omissions, and source boundaries. The lesson is not to add endless rubric dimensions; it is to align the reward with the decision.

Loss and optimization

A loss quantifies an error or undesirable objective value. An optimizer changes parameters to reduce it. A gradient describes how the loss changes with a small parameter change. You do not need to derive every optimization method to run an eval, but must understand why reducing training loss is not equivalent to improving a deployment outcome.

Overfitting means performing well on the data or patterns used for tuning while failing to generalize. A holdout helps only while it remains genuinely withheld. Training data quality, objective design, and evaluation quality are different responsibilities.

Embeddings and retrieval

An embedding maps an item to a numerical representation. Similarity search retrieves nearby representations under a chosen metric. Similarity is not truth, clinical relevance, or source authority. A semantically similar document can describe a different population or obsolete policy.

When using retrieval, retain source IDs and versions and test ranking separately from answer faithfulness. Inspect what the model actually received.

What you can defer

For a first domain eval, defer writing custom GPU kernels, building a distributed trainer, and deriving transformer architecture from first principles. For training or infrastructure roles, those subjects may become essential. Do not defer Python literacy, experimental control, data provenance, or error analysis.

Exercise: explain why a model can improve a reward score while worsening the intended task. Give one coding and one clinical-documentation example, then propose a check that detects the mismatch.