35. Choose training, prompting, or tools by the observed failure
A model-system failure should lead to a specific intervention. Prompt changes, retrieval, supervised fine-tuning, preference optimization, reinforcement learning, and infrastructure repairs change different parts of the system. Choosing among them requires a diagnosed mechanism and an acceptance test kept separate from development.
- Diagnose a mechanism before selecting an intervention.
- Distinguish SFT, preference optimization, and reward-based learning.
- Write a controlled experiment card.
Choose an intervention based on a reproduced failure mechanism. Prompting, retrieval, SFT, preference learning, rewards, and infrastructure change different parts of a system.
Introduction
The training concepts in chapter 27 describe what changes during learning. The small experiment in chapter 32 makes parameter updates visible. This chapter connects those foundations to a practical research decision, without treating every evaluation engineer as a distributed-training specialist.
| Observed mechanism | First candidate intervention | Evidence needed before escalation |
|---|---|---|
| Output schema is ambiguous | Clarify the task and validate output | Same-case comparison with schema and semantic checks |
| Required facts are absent from context | Improve eligible-source retrieval | Retrieval relevance and downstream faithfulness checks |
| Correct evidence is ignored | Compare instructions, examples, or routing | Controlled cases showing appropriate evidence use |
| Repeated domain behavior remains weak | Consider curated SFT demonstrations | Data provenance, representative failures, fresh acceptance cases |
| Competing acceptable outputs need ranking | Consider preference learning | Reliable comparisons, explicit rubric, disagreement analysis |
| Multi-step behavior has checkable outcomes | Consider reward-based optimization | Protected grader, environment fidelity, reward-gaming controls |
| A service drops requests or responses | Repair execution infrastructure | Complete cohort and fault-injection evidence |
These are candidate choices, not guaranteed fixes. A task or reference defect can make every model intervention look ineffective. Repair the measurement first when its validity is in doubt.
Understand SFT operationally
A supervised fine-tuning dataset pairs inputs with desired outputs or trajectories. Demonstrations should satisfy the intended contract and represent the relevant problem distribution. Keep provenance, annotator instructions, disagreement handling, and splits. Near-duplicate cases, leaked answers, and inconsistent demonstrations can dominate the apparent result.
A fine-tuning run also needs the base checkpoint, tokenizer, formatting template, optimizer settings, resource budget, saved outputs, and evaluation configuration. Parameter-efficient methods reduce which parameters are trained; they do not remove the need to verify the resulting behavior. A successful training job establishes that a computation completed, not that the resulting model is better.
Distinguish preference methods correctly
A conventional RLHF pipeline can include demonstrations, a preference-trained reward model, and reinforcement-learning optimization. This is one family of designs, not a mandatory sequence for every post-training method.
Direct Preference Optimization uses chosen/rejected response pairs and a preference objective relative to a reference policy, without requiring the separate reward-model fitting and online sampling stages of that conventional pipeline. It is not simply PPO under a different name. See the original DPO paper, Rafailov et al., 2023, preprint. Algorithmic simplicity does not guarantee that annotator preferences measure the deployment objective.
In a clinical-documentation comparison, a fluent answer that fills missing facts may be preferred by an inattentive rater. The reference policy should instead reward faithfulness to the source and penalize invented critical information. Preserve examples of legitimate ambiguity and adjudicated disagreement instead of declaring every comparison obvious.
Build an experiment card before spending compute
Decision: whether a bounded intervention improves the completeness task.
System: exact base checkpoint or provider identifier, prompt, tools, parser.
Failure mechanism: source date is absent, but the assistant invents it.
Intervention: demonstrations preserving unknown source dates.
Development cases: used to construct the intervention and rubric.
Acceptance cases: independently held out, with explicit provenance.
Primary check: correctly preserve present and absent source dates.
Critical check: no fabricated source dates in the acceptance cases.
Controls: valid dates, absent dates, conflicting versions, malformed input.
Budget: approved run limits and stop conditions.
Report: complete denominator, matched differences, failures, uncertainty.
This is an invented experiment specification, not a completed fine-tuning study. Thresholds and sample size must be set for the actual decision; the card supplies no clinical acceptance threshold.
Use coding agents to accelerate implementation
Ask an agent to build an adapter, dataset validator, split checker, or report generator under a precise contract. Require the exact executed command, exit status, artifacts, and error tests. Inspect the data flow and diff. Keep hidden acceptance references outside the agent’s accessible workspace when the experimental design requires that separation.
For a training role, add hands-on framework work, checkpoint recovery, memory and throughput profiling, distributed failure handling, and algorithm-specific mathematics. For an evaluation role, first produce a valid measured comparison and a defensible failure analysis. The depth follows the role and demonstrated bottleneck.
Checkpoint: choose an intervention for a missing-context failure, an invalid-output failure, and a dropped-response failure. State what each intervention changes, the control that could falsify its benefit, and the evidence required before expanding the work.