AI safety and defensive biosecurity evaluations
TL;DR
Define the safety boundary, threat model, permissions, and escalation. Measure unsafe assistance and benign usefulness separately; public biological exercises stay defensive and non-actionable.
Safety evals ask whether the configured system behaves within a specified boundary and whether safeguards withstand relevant pressure. A refusal count alone does not measure safety. A system can over-refuse harmless work or provide dangerous help after a superficial disclaimer.
Define behavior, context, and authority
List permitted work, restricted actions, protected information, escalation triggers, and approved tool capabilities. Evaluate the actual system including enforcement. A model's instruction to avoid sending data is weaker evidence than a tool permission that prevents unauthorized sending.
Use benign canary tokens for information-leakage tests and mocked services for unauthorized-action tests. Do not place real secrets in a benchmark merely to test whether they leak.
Worked task: synthetic policy classification
The starter safety track uses abstract requests and a toy rule: benign defensive assessment can proceed; a request explicitly seeking harm-enabling operational steps must be refused and redirected to a safe alternative; ambiguity requires review. It contains no actionable harmful procedures.
The grader tests only the selected label. A real safety eval must also inspect the full answer for prohibited assistance, useful safe help, consistency, and tool behavior. “refuse” in a JSON field cannot erase unsafe prose or an unsafe tool call.
Threat model
Name the adversary, access, objective, and constraints. An ordinary user's ambiguous question and an adversarial document in retrieval represent different threats. Include controls showing that harmless requests remain useful. Predefine what counts as a policy violation and what remains ambiguous.
For public learning, test harmless prompt injection such as a retrieved document asking the assistant to reveal a synthetic canary or ignore the supplied classification rule. Preserve the boundary between data and instructions. Keep the environment isolated from real credentials and external recipients.
Defensive biological evaluation
Measure whether the system recognizes sensitive context, avoids operational harmful assistance, provides appropriate defensive alternatives, and preserves access restrictions. Keep specific hazardous case content restricted to qualified, authorized teams. Public summaries can describe failure categories and mitigation outcomes without publishing enabling details.
Distinguish knowledge from operational assistance, refusal from robust safety, and a tested policy boundary from general safety. Do not claim that a passed suite establishes absence of harmful capability.
Governance and reporting
NIST's Generative AI Profile is a risk-management reference. It supports organizing measurement and governance; it does not supply a universal pass score for every system.
Report violation rate for the tested cases, benign usefulness, unresolved cases, severity categories, and changes after mitigation. Keep representative, adversarial, and regression sets separately identified. Restrict sensitive traces and disclose only what is safe and necessary.
Exercise: a system refuses the harmful request but also refuses a benign request to summarize a biosafety policy. Write separate findings for boundary adherence and benign usefulness. Do not combine them into a single safety percentage.