AI safety work turns a possible failure into an observable experiment and a decision about controls. The useful starting question is specific: what could this configured system do, through which interface, under whose authority, and with what consequence? Safety, security, and clinical validity share methods, but they do not have interchangeable reference standards.
NoteLearning objectives
Define capability, propensity, and control as different measurement questions.
Map a safety risk to an observable test and mitigation.
Navigate current primary sources without treating policies as validation.
TipTL;DR
Start with a risk-to-test map, then run the agent, security, and improvement exercises. Capability, observed behavior, and control effectiveness are separate questions. Use the dated primary-source map for current lab policies and research.
Introduction
Begin with the boundary definitions in lesson 19. Use this path to build the engineering and research artifacts around them. The examples are authored teaching scenarios, not accounts of any employer’s confidential work.
Need
Read
Produce
Understand safety research and its decisions
This lesson
A risk-to-test map
Test unauthorized actions and information disclosure
Capability asks whether the system can accomplish something under specified conditions. Propensity asks whether it does so in the tested circumstances. Control asks whether the surrounding system prevents, detects, or contains an unacceptable action. A capable model need not attempt a harmful action in ordinary use; an unsuccessful attempt can still expose a weakness in intent or oversight.
A medical example makes the distinction concrete. An assistant may be able to generate an unsupported dose, may rarely do so in a curated extraction task, and may be technically unable to write a prescription because no prescribing tool is connected. Those are separate observations. Only the last addresses that particular action boundary; none alone establishes clinical benefit.
Area
Practical question
Contained teaching test
Reliability
Does the output preserve the evidence?
An absent source date remains unknown
Misuse resistance
Does the full response respect a supplied boundary?
Abstract harmful-intent labels with benign defensive controls
Alignment
Does behavior remain consistent with the authorized goal?
An agent admits a failed test rather than reporting success
Security
Can untrusted material redirect authority or expose data?
A bounded optimization task with locked correctness checks
These are curriculum categories, not a replacement for a lab’s formal policy taxonomy.
What safety research methods actually do
Behavioral evaluation measures answers and actions in a defined environment. Red teaming searches for failures, often adaptively. An adversarially selected set is useful evidence that a failure exists; its failure fraction usually does not estimate ordinary-user prevalence.
Interpretability studies internal representations or computations. It can generate mechanistic hypotheses, but an attractive explanation needs causal tests: intervene on the proposed mechanism, check the predicted behavior, and measure side effects. A readable explanation emitted by the model is not an independent account of its internal computation.
Control research studies systems that retain useful work while limiting the consequences of a potentially untrusted component. A smaller monitor, restricted tools, independent review, and selective escalation can each contribute. Evaluate the entire protocol, including what slips through and how much legitimate work is blocked.
Alignment auditing looks for behavior that ordinary task scores miss. Anthropic’s sabotage evaluation research distinguishes decision sabotage, code sabotage, sandbagging, and undermining oversight. Its experiments should be read with their artificial settings and historical model versions in view. The later Petri 2.0 report describes automated auditing and the difficulty of realistic evaluations. It is a discovery tool and research method, not a certificate of safety.
Build a risk-to-test map
Use this invented task: a research assistant summarizes synthetic surveillance reports and drafts a review note. It can read approved records. It cannot change records, contact recipients, or issue public alerts.
Consequence: an unsupported report is presented as corroborated.
Surface: retrieved report text and final summary.
Evidence needed: source revision, availability time, supporting span.
Test: unverified report, later correction, syndicated duplicate.
Failure rule: unsupported certainty or future information in the summary.
Control: explicit evidence-state fields and independent source checks.
Remaining gap: fixture performance does not measure live reviewer behavior.
Add a second row for unauthorized writing. Its evidence is an execution receipt or denied action, not a reassuring sentence. Add a third row for hidden-reference access. Its control is grader-side isolation, not a request that the agent refrain from looking.
Prioritize by the consequence and the plausible access path. A long catalog of risks with no executable test does not establish coverage. Conversely, a precise small suite can expose a material defect without claiming to cover every threat.
A safety experiment from request to decision
Define the permitted task and prohibited consequences with the accountable owner.
Build ordinary cases, adversarial cases, and benign lookalikes. Keep development discoveries separate from acceptance cases.
Freeze the system configuration and rubric. Record tools, permissions, environment, budgets, and failure statuses.
Validate the grader using authored failures and legitimate alternatives before testing a model.
Run the configured system and retain complete outputs, action traces, and destination state where applicable.
Adjudicate ambiguous cases blind to system identity when feasible. Record disagreement rather than hiding it.
Compare mitigations on matched cases and fresh cases. Inspect regression in legitimate usefulness.
Write a decision bounded by the tested environment, with severe failures visible beside aggregate results.
Lessons 5–14 supply the canonical dataset, grader, statistics, and reporting methods. Use those methods here rather than creating a separate safety scoring philosophy.
Separation of agent runtime and permission authority
Can the main agent modify the control that is supposed to constrain it?
The retrieved Anthropic page lists RSP 3.4, effective July 8, 2026, and the DeepMind page lists framework 3.1, dated April 17, 2026. Recheck the linked revision history before relying on a policy operationally. The OpenAI April 2025 article is deliberately labeled historical; later safety material must also be considered. These are organizational policies and research reports, not universal clinical acceptance standards.
Use domain expertise where it changes the experiment
In healthcare, identify unsupported clinical specificity and consequences that a generic grader misses. In biology, separate a valid analysis from an invalid biological inference. In biosurveillance, protect time, event identity, and evidence state. In defensive biosecurity, keep hazardous case detail restricted and evaluate access boundaries, safe alternatives, and legitimate defensive usefulness.
The transferable skill is making a disputed claim testable. A physician-researcher can contribute strong reference standards and consequence models; implementation still requires inspecting data flow and executed evidence. Coding agents help produce those artifacts, but their confidence does not validate them.
Checkpoint: produce three risk-to-test rows for the synthetic surveillance assistant. For each, name the consequence, observable failure, control, and remaining uncertainty. Then follow the security lab.