39. Recursive self-improvement and AI research evaluations
Recursive self-improvement concerns systems improving the systems or processes that produce their future capabilities. The term is used for materially different feedback loops. An evaluation must name what changes, who supplies the objective, what evidence validates the change, and how much of the loop operates without human intervention.
- Distinguish answer refinement, harness changes, training, and autonomous R&D.
- Run an integrity-checked comparison with critical failure gates.
- Bound self-improvement claims by independent evidence and selection history.
Name what improves and who controls the loop. AI-assisted research, harness optimization, parameter updates, and autonomous successor development require different evidence. Run the frozen comparison and reject incomplete cohorts, regressions, critical failures, or altered references.
Introduction
Use the training distinctions in lessons 27 and 35 before interpreting a self-improvement claim. Editing instructions, changing runtime code, training weights, and redesigning a research program modify different objects. A conversation that revises an answer does not by itself update model weights.
Distinguish the loops
| Loop | Object changed | Meaningful evidence |
|---|---|---|
| Answer refinement | Current answer or plan | Better result on independent task checks |
| Harness improvement | Prompt, tools, context selection, runtime | Fresh task gains with regression and security checks |
| Parameter update | Model weights through training | Held-out improvement under a controlled comparison |
| AI-assisted R&D | Experiments, analysis, engineering | Reproduced findings and researcher intervention records |
| Autonomous successor development | Design, training, evaluation of a successor | End-to-end completion and independently verified successor capability |
| Recursive improvement of the improvement process | The next cycle becomes more effective | Repeated controlled cycles that improve on new tasks |
Definitions vary across research. State the operational definition in the report rather than treating every improvement as the same phenomenon. More code, more experiments, or a better public benchmark score does not automatically establish better general capability or a faster recursive process.
What the primary sources establish
Anthropic’s When AI builds itself reports AI contributions to its engineering and research and explicitly distinguishes these from fully autonomous successor development. It identifies continuing gaps in goal selection. Its internal productivity measures and task results are company-reported evidence with measurement limitations, not an independently verified demonstration of inevitable runaway improvement.
The more recent automated alignment-research report, dated August 28, 2026, describes improvements on alignment-related benchmarks. The report cautions that those evaluations are proxies for real-world misalignment and that persistence after extensive subsequent RL was not tested. Read the experiment’s actual scope before generalizing to long-term safety.
METR’s RE-Bench provides ML research-engineering environments for AI R&D evaluations. Completing such a task is evidence about that environment and budget. It does not by itself demonstrate autonomous selection of a valuable research direction, frontier training capability, or a full successor-development loop. Respect benchmark access and solution restrictions; do not expose protected solutions to the tested agent.
Calibrate the claim
Demonstrated means the reported experiment produced the stated result under its documented conditions. Theoretical means the mechanism is plausible but the full claim has not been demonstrated. Beyond current evidence means the available source does not establish the proposed capability or extrapolation.
An automated program can demonstrably select a better candidate in a contained experiment. Whether that selection improves generalization is a separate empirical question. Whether repeated cycles yield unbounded acceleration is a much stronger claim. A chart extrapolation or an executive forecast is not that demonstration.
Keep this vocabulary in research notes:
Object changed: harness code, not model weights.
Objective source: human-specified task and grader.
Autonomy: candidate generation automated; acceptance independently reviewed.
Demonstrated result: complete held-out comparison in the stated environment.
Unresolved: transfer, adaptive overfitting, monitoring gaps, longer horizons.
A controlled research task
Use the small classifier in lesson 32. Ask an agent to reduce training runtime while preserving independently checked output behavior. Keep hardware, data, precision, initialization, and correctness criteria fixed. Record warm-up policy and repeated timings; do not call one favorable run a speedup.
Give the agent development cases and an approved resource budget. Keep acceptance labels and the grader outside its writable environment. Review whether a purported optimization skipped cases, changed precision, reduced training steps, or weakened tests. Those changes can be legitimate experiments, but they cannot be reported as the same computation without qualification.
For a biosurveillance research task, permit a change to evidence retrieval or event linking. Preserve the temporal replay and hidden references. Evaluate whether fewer duplicate alerts came at the cost of missed distinct events. An optimization of a proxy can damage the mission outcome.
Run the improvement gate
python3 foundations/improvement_gate.pyThe authored baseline has two correct outcomes among three cases. Both candidates have three correct labels, but one has a critical failure. The first is CANDIDATE_FOR_REVIEW; the second is HOLD. Neither status deploys anything. A candidate still needs independent review and a decision appropriate to the real use.
from __future__ import annotations
import hashlib
import json
from dataclasses import dataclass
@dataclass(frozen=True)
class Outcome:
case_id: str
correct: bool
critical_failure: bool
def fingerprint(references: dict[str, str]) -> str:
if not references or any(not k.strip() or not v.strip() for k, v in references.items()):
raise ValueError("nonempty reference IDs and labels required")
payload = json.dumps(references, sort_keys=True, separators=(",", ":")).encode()
return hashlib.sha256(payload).hexdigest()
def validate(outcomes: list[Outcome], ids: set[str]) -> dict[str, Outcome]:
if len(outcomes) != len(ids) or {o.case_id for o in outcomes} != ids:
raise ValueError("missing, extra, or duplicate outcomes")
if any(type(o.correct) is not bool or type(o.critical_failure) is not bool for o in outcomes):
raise ValueError("grader outcomes must be booleans")
return {o.case_id: o for o in outcomes}
def compare(references: dict[str, str], expected_hash: str,
baseline: list[Outcome], candidate: list[Outcome]) -> dict[str, object]:
if fingerprint(references) != expected_hash:
raise ValueError("protected references changed")
old = validate(baseline, set(references))
new = validate(candidate, set(references))
gained = sorted(k for k in old if not old[k].correct and new[k].correct)
regressed = sorted(k for k in old if old[k].correct and not new[k].correct)
critical = sorted(k for k in new if new[k].critical_failure)
# An invented conservative teaching gate, not a deployment threshold.
eligible = bool(gained) and not regressed and not critical
return {"status": "CANDIDATE_FOR_REVIEW" if eligible else "HOLD",
"cases": len(old), "baseline_correct": sum(o.correct for o in old.values()),
"candidate_correct": sum(o.correct for o in new.values()),
"gained": gained, "regressed": regressed, "critical": critical,
"reference_hash": expected_hash}
def main() -> None:
references = {"a": "review", "b": "complete", "c": "review"}
baseline = [Outcome("a", True, False), Outcome("b", True, False), Outcome("c", False, False)]
improved = [Outcome("a", True, False), Outcome("b", True, False), Outcome("c", True, False)]
unsafe = [Outcome("a", True, False), Outcome("b", True, False), Outcome("c", True, True)]
for name, candidate in [("improved", improved), ("unsafe", unsafe)]:
print(json.dumps({"fixture": name, **compare(references, fingerprint(references), baseline, candidate)},
sort_keys=True))
if __name__ == "__main__":
main()The input outcomes are trusted grader results, not assertions supplied by the candidate. The reference hash is checked against a value held by the supervisor. In this offline demonstration, everything is readable in one folder; the hash illustrates integrity comparison and provides no secrecy, authentication, or access control. A real evaluator must protect the expected hash and reference store independently.
The gate rejects missing cases, extra cases, duplicate cases, changed references, and malformed outcomes. It requires at least one gain, no regressions, and no critical failures. This conservative rule is an invented teaching policy. Statistical decisions and acceptable tradeoffs require the methods in lesson 12 and the actual task’s consequence model.
Protect the improvement loop from reward gaming
Reward hacking occurs when the measured objective rises without the intended outcome improving. It can come from ordinary optimization pressure without proving deliberate deception. A candidate that changes the grader, copies hidden labels, omits difficult cases, or produces a score instead of an artifact defeats the experiment.
Protect independent acceptance evidence, inspect artifacts, and use task variants the optimizer has not seen. A public holdout repeatedly used to choose changes gradually becomes development data. Even an inaccessible holdout can become adaptively overfit when detailed feedback from many trials drives the next candidate. Limit feedback and budget, retain a final independent set, and document selection history.
Monitor the monitor. Seed harmless known failures and measure whether the oversight process catches them. Evaluate missed violations, false alarms, review workload, and whether the agent can modify its own logs or change the monitor configuration. A monitor’s verbal assurance is not an audit receipt.
A practical AI R&D evaluation specification
| Requirement | Example evidence |
|---|---|
| Research objective and scope | Fixed training optimization task; no network or deployment |
| Environment | Recorded versions, hardware, data hashes, precision |
| Baselines | Original implementation and qualified human comparator where justified |
| Budget | Wall time, calls, compute, attempts, intervention policy |
| Correctness | Independent numerical and behavioral checks |
| Improvement | Repeated timing or fresh task result, not a self-reported score |
| Integrity | Grader and reference hashes, denied writes, preserved artifacts |
| Autonomy | Human hints, resets, decisions, and takeovers recorded |
| Safety | Permissions, cancellation, restricted outputs, monitor coverage |
| Conclusion | Scope-limited finding, regressions, uncertainty, next test |
This work joins evaluation engineering with research judgment. A researcher chooses the construct and experimental design; an engineer makes execution reliable; a reviewer tests whether the claimed advance survives independent checks.
Checkpoint: classify a prompt rewrite, a training run, and an automated research session by object changed and autonomy. Run the gate, then remove one candidate outcome in a local copy and confirm it raises instead of reporting a favorable partial score.