Bryan Tegomoh / Learn

Learn to evaluate AI.
Make the evidence count.

A practical field manual for a physician-scientist entering AI evaluation. From your first reproducible test to coding-agent reviews and domain-specific research.

Engineering judgmentScientific measurementDefensive evaluation

Your expertise is the starting point.

Bring the rigor of medicine and epidemiology. Learn the system, define the reference standard, inspect failures, and use coding agents to accelerate the work you can verify.

AI tools accelerate implementation. Staff-level judgment still comes from practice. This manual gives you a route to demonstrable competence, with explicit checkpoints.

Choose your route

First-day essentialsUnderstand the loop. Run the kit. Review a repair.Health & scientific evaluationMedicine, biology, biosurveillance, and AI safety.What employers actually ask forOfficial postings translated into practical learning goals.

Learn by running an evaluation.

Local Python kit, synthetic cases, strict graders, tests, and templates. No model API calls.

Download the starter kit

The complete manual

Read in order or go directly to the problem you need to solve.

01

Start here: what this job demands

Your job is to turn a claim into a defensible experiment. Start with one task, one valid grader, one inspected failure, and one report. AI tools accelerate implementation; staff-level judgment requires practice.

02

The minimum engineering you need to understand

Understand functions, types, state, side effects, errors, and tests well enough to inspect generated code. Focus on Python and the invariants a change must preserve.

03

Set up a workspace and inspect a repository

Confirm the folder, instructions, dependencies, and baseline before editing. Record versions and inspect both the diff and untracked files so the tested revision is known.

04

Use Codex, Claude Code, and Cursor effectively

Give coding agents a bounded contract and demand executed evidence. Use them for implementation and investigation while you own the objective, reference, permissions, and conclusion.

05

Design an evaluation before building it

Define the decision, task, population, inputs, outputs, reference, and failure criteria before building. Match the strength of your conclusion to the construct actually tested.

06

Build datasets and reference standards

Build versioned cases with provenance and defensible labels. Keep answers out of model-visible inputs and separate development from a genuinely withheld acceptance set.

07

Grade outcomes, behavior, and evidence

Use deterministic graders when sufficient and calibrated expert or model review when needed. Test the grader with valid alternatives, misleading outputs, and critical failures.

08

Run your first complete evaluation

Run the local kit, inspect every score, and demonstrate rejection of invalid or incomplete responses. Fixture scores verify the apparatus, not a model or clinical workflow.

09

Review a coding agent end to end

Review the full agent interaction, exact patch, resulting behavior, and verification. Tie findings to a violated requirement and reproducible evidence, not personal style preferences.

10

Debug systematically and assess architecture

Reproduce the failure, test competing hypotheses, and repair the smallest verified cause. Judge architecture by state ownership, failure handling, and actual requirements.

11

Evaluate agents, tools, retrieval, and memory

Evaluate the configured model, tools, retrieval, memory, and permissions together. Separate final outcomes from trajectories and inspect whether retrieved evidence truly supports claims.

12

Statistics, uncertainty, and comparisons

Name the denominator, sampling unit, dependence, and uncertainty before interpreting a score. Compare matched cases and keep critical failures visible beside averages.

13

Diagnose failures and improve the system

Classify failures by mechanism so the next experiment has a purpose. Repair defective tasks or graders transparently and preserve a fresh acceptance set after tuning.

14

Write reviews, calibration notes, and decision reports

Write the result first, then the requirement, evidence, consequence, and uncertainty. A good report supports a specific decision without overstating what was tested.

15

Health systems and public health evaluations

Distinguish administrative, operational, population, and patient-facing tasks. Evaluate routing and workload as well as content, with explicit references and governance.

16

Medicine: clinical reasoning and documentation

Separate extraction, documentation, clinical reasoning, and treatment tasks. Preserve unknowns, negation, entity identity, and time; software checks do not establish clinical benefit.

17

Biology and life sciences evaluations

Test scientific evidence handling, units, cohorts, and reproducibility. Working code can still support an invalid scientific conclusion; keep biological teaching cases benign.

18

Biosurveillance and epidemic intelligence evaluations

Preserve information available at the decision time and define event-level targets. Evaluate delay, alert burden, uncertainty, and revisions instead of relying on a single retrospective score.

19

AI safety and defensive biosecurity evaluations

Define the safety boundary, threat model, permissions, and escalation. Measure unsafe assistance and benign usefulness separately; public biological exercises stay defensive and non-actionable.

20

Privacy, security, and operational discipline

Classify data, enforce least privilege, protect logs, and establish budgets before paid runs. Use synthetic cases for learning and institutionally authorized arrangements for real sensitive data.

21

From an eval notebook to a reliable evaluation service

Version cases, control execution, separate adapters from graders, and preserve complete reports. Expose failed and cancelled work rather than silently shrinking the denominator.

22

An apprenticeship plan and portfolio

Build a portfolio through increasingly demanding artifacts, not self-awarded titles. A small reproducible project with honest limits is stronger than an impressive but unverifiable demo.

23

Calibration interview and readiness assessment

Practice explaining the system, reference, grader, failures, and limits. Staff-level coding evaluation needs demonstrated engineering judgment beyond reading this course.

24

Templates, prompts, and reusable artifacts

Use task, review, failure, and decision templates to preserve inspectable evidence. Unknowns stay unknown; completed fields should represent work that actually happened.

25

Glossary and primary-source reading map

Use the glossary while building artifacts and consult primary documentation for changing interfaces. Sources guide learning; teaching exercises are not published research results.

26

What frontier labs are hiring for now

Selected live official postings ask for engineering execution, valid measurement, diagnosis, and communication. Some welcome self-taught AI power users, but none of these examples makes coding competence optional.

27

ML and post-training concepts in plain language

Understand training, inference, SFT, preference learning, rewards, and RL in terms of what changes and what is optimized. Better reward scores can still conceal worse real behavior.

28

Lab walkthroughs with solutions and review notes

Work through repairs, strict scoring, missingness, rare-event metrics, temporal leakage, and judge checks. Solutions explain the mechanism and state exactly what each exercise proves.

29

Advanced engineering practice without unnecessary theory

Deepen APIs, joins, concurrency, containers, testing, and ML frameworks when your role requires them. Reading a concept is preparation; a tested artifact is evidence of competence.

30

Your focused syllabus and ways of thinking

Learn the common evaluation loop first, then specialize by the role you want. Progress when you can produce, explain, test, and bound each artifact without outsourcing your judgment.