Learn to evaluate AI.
Make the evidence count.
A practical field manual for a physician-scientist entering AI evaluation. From your first reproducible test to coding-agent reviews and domain-specific research.
Your expertise is the starting point.
Bring the rigor of medicine and epidemiology. Learn the system, define the reference standard, inspect failures, and use coding agents to accelerate the work you can verify.
AI tools accelerate implementation. Staff-level judgment still comes from practice. This manual gives you a route to demonstrable competence, with explicit checkpoints.
Choose your route
First-day essentialsUnderstand the loop. Run the kit. Review a repair.Health & scientific evaluationMedicine, biology, biosurveillance, and AI safety.What employers actually ask forOfficial postings translated into practical learning goals.Learn by running an evaluation.
Local Python kit, synthetic cases, strict graders, tests, and templates. No model API calls.
The complete manual
Read in order or go directly to the problem you need to solve.
Start here: what this job demands
Your job is to turn a claim into a defensible experiment. Start with one task, one valid grader, one inspected failure, and one report. AI tools accelerate implementation; staff-level judgment requires practice.
The minimum engineering you need to understand
Understand functions, types, state, side effects, errors, and tests well enough to inspect generated code. Focus on Python and the invariants a change must preserve.
Set up a workspace and inspect a repository
Confirm the folder, instructions, dependencies, and baseline before editing. Record versions and inspect both the diff and untracked files so the tested revision is known.
Use Codex, Claude Code, and Cursor effectively
Give coding agents a bounded contract and demand executed evidence. Use them for implementation and investigation while you own the objective, reference, permissions, and conclusion.
Design an evaluation before building it
Define the decision, task, population, inputs, outputs, reference, and failure criteria before building. Match the strength of your conclusion to the construct actually tested.
Build datasets and reference standards
Build versioned cases with provenance and defensible labels. Keep answers out of model-visible inputs and separate development from a genuinely withheld acceptance set.
Grade outcomes, behavior, and evidence
Use deterministic graders when sufficient and calibrated expert or model review when needed. Test the grader with valid alternatives, misleading outputs, and critical failures.
Run your first complete evaluation
Run the local kit, inspect every score, and demonstrate rejection of invalid or incomplete responses. Fixture scores verify the apparatus, not a model or clinical workflow.
Review a coding agent end to end
Review the full agent interaction, exact patch, resulting behavior, and verification. Tie findings to a violated requirement and reproducible evidence, not personal style preferences.
Debug systematically and assess architecture
Reproduce the failure, test competing hypotheses, and repair the smallest verified cause. Judge architecture by state ownership, failure handling, and actual requirements.
Evaluate agents, tools, retrieval, and memory
Evaluate the configured model, tools, retrieval, memory, and permissions together. Separate final outcomes from trajectories and inspect whether retrieved evidence truly supports claims.
Statistics, uncertainty, and comparisons
Name the denominator, sampling unit, dependence, and uncertainty before interpreting a score. Compare matched cases and keep critical failures visible beside averages.
Diagnose failures and improve the system
Classify failures by mechanism so the next experiment has a purpose. Repair defective tasks or graders transparently and preserve a fresh acceptance set after tuning.
Write reviews, calibration notes, and decision reports
Write the result first, then the requirement, evidence, consequence, and uncertainty. A good report supports a specific decision without overstating what was tested.
Health systems and public health evaluations
Distinguish administrative, operational, population, and patient-facing tasks. Evaluate routing and workload as well as content, with explicit references and governance.
Medicine: clinical reasoning and documentation
Separate extraction, documentation, clinical reasoning, and treatment tasks. Preserve unknowns, negation, entity identity, and time; software checks do not establish clinical benefit.
Biology and life sciences evaluations
Test scientific evidence handling, units, cohorts, and reproducibility. Working code can still support an invalid scientific conclusion; keep biological teaching cases benign.
Biosurveillance and epidemic intelligence evaluations
Preserve information available at the decision time and define event-level targets. Evaluate delay, alert burden, uncertainty, and revisions instead of relying on a single retrospective score.
AI safety and defensive biosecurity evaluations
Define the safety boundary, threat model, permissions, and escalation. Measure unsafe assistance and benign usefulness separately; public biological exercises stay defensive and non-actionable.
Privacy, security, and operational discipline
Classify data, enforce least privilege, protect logs, and establish budgets before paid runs. Use synthetic cases for learning and institutionally authorized arrangements for real sensitive data.
From an eval notebook to a reliable evaluation service
Version cases, control execution, separate adapters from graders, and preserve complete reports. Expose failed and cancelled work rather than silently shrinking the denominator.
An apprenticeship plan and portfolio
Build a portfolio through increasingly demanding artifacts, not self-awarded titles. A small reproducible project with honest limits is stronger than an impressive but unverifiable demo.
Calibration interview and readiness assessment
Practice explaining the system, reference, grader, failures, and limits. Staff-level coding evaluation needs demonstrated engineering judgment beyond reading this course.
Templates, prompts, and reusable artifacts
Use task, review, failure, and decision templates to preserve inspectable evidence. Unknowns stay unknown; completed fields should represent work that actually happened.
Glossary and primary-source reading map
Use the glossary while building artifacts and consult primary documentation for changing interfaces. Sources guide learning; teaching exercises are not published research results.
What frontier labs are hiring for now
Selected live official postings ask for engineering execution, valid measurement, diagnosis, and communication. Some welcome self-taught AI power users, but none of these examples makes coding competence optional.
ML and post-training concepts in plain language
Understand training, inference, SFT, preference learning, rewards, and RL in terms of what changes and what is optimized. Better reward scores can still conceal worse real behavior.
Lab walkthroughs with solutions and review notes
Work through repairs, strict scoring, missingness, rare-event metrics, temporal leakage, and judge checks. Solutions explain the mechanism and state exactly what each exercise proves.
Advanced engineering practice without unnecessary theory
Deepen APIs, joins, concurrency, containers, testing, and ML frameworks when your role requires them. Reading a concept is preparation; a tested artifact is evidence of competence.
Your focused syllabus and ways of thinking
Learn the common evaluation loop first, then specialize by the role you want. Progress when you can produce, explain, test, and bound each artifact without outsourcing your judgment.