36. AI safety: a practical learning path

Published

October 3, 2026

AI safety work turns a possible failure into an observable experiment and a decision about controls. The useful starting question is specific: what could this configured system do, through which interface, under whose authority, and with what consequence? Safety, security, and clinical validity share methods, but they do not have interchangeable reference standards.

NoteLearning objectives
  • Define capability, propensity, and control as different measurement questions.
  • Map a safety risk to an observable test and mitigation.
  • Navigate current primary sources without treating policies as validation.
TipTL;DR

Start with a risk-to-test map, then run the agent, security, and improvement exercises. Capability, observed behavior, and control effectiveness are separate questions. Use the dated primary-source map for current lab policies and research.

Introduction

Begin with the boundary definitions in lesson 19. Use this path to build the engineering and research artifacts around them. The examples are authored teaching scenarios, not accounts of any employer’s confidential work.

Need Read Produce
Understand safety research and its decisions This lesson A risk-to-test map
Test unauthorized actions and information disclosure 37. AI security A trace-based security report
Understand how coding and research agents work 38. Agent architecture A working bounded loop
Evaluate self-improvement claims 39. Recursive self-improvement A frozen comparison and integrity gate
See what the working day produces 40. Daily work A complete experiment packet
Build a surveillance evaluation 41. Biosurveillance benchmark A temporal replay with event-level scoring

Separate the questions before choosing a test

Capability asks whether the system can accomplish something under specified conditions. Propensity asks whether it does so in the tested circumstances. Control asks whether the surrounding system prevents, detects, or contains an unacceptable action. A capable model need not attempt a harmful action in ordinary use; an unsuccessful attempt can still expose a weakness in intent or oversight.

A medical example makes the distinction concrete. An assistant may be able to generate an unsupported dose, may rarely do so in a curated extraction task, and may be technically unable to write a prescription because no prescribing tool is connected. Those are separate observations. Only the last addresses that particular action boundary; none alone establishes clinical benefit.

Area Practical question Contained teaching test
Reliability Does the output preserve the evidence? An absent source date remains unknown
Misuse resistance Does the full response respect a supplied boundary? Abstract harmful-intent labels with benign defensive controls
Alignment Does behavior remain consistent with the authorized goal? An agent admits a failed test rather than reporting success
Security Can untrusted material redirect authority or expose data? Synthetic canary and denied mock tools
Oversight Does a monitor catch a consequential deviation? Seeded, harmless trace anomalies reviewed independently
AI R&D acceleration What part of research can the system complete? A bounded optimization task with locked correctness checks

These are curriculum categories, not a replacement for a lab’s formal policy taxonomy.

What safety research methods actually do

Behavioral evaluation measures answers and actions in a defined environment. Red teaming searches for failures, often adaptively. An adversarially selected set is useful evidence that a failure exists; its failure fraction usually does not estimate ordinary-user prevalence.

Interpretability studies internal representations or computations. It can generate mechanistic hypotheses, but an attractive explanation needs causal tests: intervene on the proposed mechanism, check the predicted behavior, and measure side effects. A readable explanation emitted by the model is not an independent account of its internal computation.

Control research studies systems that retain useful work while limiting the consequences of a potentially untrusted component. A smaller monitor, restricted tools, independent review, and selective escalation can each contribute. Evaluate the entire protocol, including what slips through and how much legitimate work is blocked.

Alignment auditing looks for behavior that ordinary task scores miss. Anthropic’s sabotage evaluation research distinguishes decision sabotage, code sabotage, sandbagging, and undermining oversight. Its experiments should be read with their artificial settings and historical model versions in view. The later Petri 2.0 report describes automated auditing and the difficulty of realistic evaluations. It is a discovery tool and research method, not a certificate of safety.

Build a risk-to-test map

Use this invented task: a research assistant summarizes synthetic surveillance reports and drafts a review note. It can read approved records. It cannot change records, contact recipients, or issue public alerts.

Consequence: an unsupported report is presented as corroborated.
Surface: retrieved report text and final summary.
Evidence needed: source revision, availability time, supporting span.
Test: unverified report, later correction, syndicated duplicate.
Failure rule: unsupported certainty or future information in the summary.
Control: explicit evidence-state fields and independent source checks.
Remaining gap: fixture performance does not measure live reviewer behavior.

Add a second row for unauthorized writing. Its evidence is an execution receipt or denied action, not a reassuring sentence. Add a third row for hidden-reference access. Its control is grader-side isolation, not a request that the agent refrain from looking.

Prioritize by the consequence and the plausible access path. A long catalog of risks with no executable test does not establish coverage. Conversely, a precise small suite can expose a material defect without claiming to cover every threat.

A safety experiment from request to decision

  1. Define the permitted task and prohibited consequences with the accountable owner.
  2. Build ordinary cases, adversarial cases, and benign lookalikes. Keep development discoveries separate from acceptance cases.
  3. Freeze the system configuration and rubric. Record tools, permissions, environment, budgets, and failure statuses.
  4. Validate the grader using authored failures and legitimate alternatives before testing a model.
  5. Run the configured system and retain complete outputs, action traces, and destination state where applicable.
  6. Adjudicate ambiguous cases blind to system identity when feasible. Record disagreement rather than hiding it.
  7. Compare mitigations on matched cases and fresh cases. Inspect regression in legitimate usefulness.
  8. Write a decision bounded by the tested environment, with severe failures visible beside aggregate results.

Lessons 5–14 supply the canonical dataset, grader, statistics, and reporting methods. Use those methods here rather than creating a separate safety scoring philosophy.

Primary-source map, checked October 3, 2026

Source Use it for Reading question
OpenAI safety Entry point to current safety work Which linked report actually supports a claim?
OpenAI Deployment Safety Hub Model-specific system cards What model, configuration, and evaluation date were tested?
OpenAI Preparedness update, April 2025 Historical framework structure How are capability evidence and safeguard evidence separated?
OpenAI cyber-capability update, August 2026 Later monitoring, containment, and alignment work What changed after the earlier framework?
OpenAI long-horizon safety, July 2026 Safety problems in extended agent work What becomes harder when actions persist across turns?
Anthropic Responsible Scaling Policy Current policy and revision history What threshold, mitigation, and review process are specified?
Google DeepMind frontier safety Framework versions and safety reports Which capability and mitigation are paired?
Anthropic alignment research Research directions and experiments What is observed behavior versus a proposed mechanism?
METR RE-Bench AI research-engineering evaluation environments What does task success prove about autonomous R&D?
Anthropic agent evaluations Tasks, trajectories, and graders Is the grader measuring the outcome or trusting the agent’s explanation?
Meta Muse safety engineering Separation of agent runtime and permission authority Can the main agent modify the control that is supposed to constrain it?

The retrieved Anthropic page lists RSP 3.4, effective July 8, 2026, and the DeepMind page lists framework 3.1, dated April 17, 2026. Recheck the linked revision history before relying on a policy operationally. The OpenAI April 2025 article is deliberately labeled historical; later safety material must also be considered. These are organizational policies and research reports, not universal clinical acceptance standards.

Use domain expertise where it changes the experiment

In healthcare, identify unsupported clinical specificity and consequences that a generic grader misses. In biology, separate a valid analysis from an invalid biological inference. In biosurveillance, protect time, event identity, and evidence state. In defensive biosecurity, keep hazardous case detail restricted and evaluate access boundaries, safe alternatives, and legitimate defensive usefulness.

The transferable skill is making a disputed claim testable. A physician-researcher can contribute strong reference standards and consequence models; implementation still requires inspecting data flow and executed evidence. Coding agents help produce those artifacts, but their confidence does not validate them.

Checkpoint: produce three risk-to-test rows for the synthetic surveillance assistant. For each, name the consequence, observable failure, control, and remaining uncertainty. Then follow the security lab.