Your focused learning syllabus
Learn the common evaluation loop, then specialize. Each stage ends with an artifact you can inspect and explain. Reading progress is a convenience, not a competence score.
TL;DR
Today: understand the system, define a task, run the kit, inspect a failure, and review a coding repair. Next: build a domain reference standard and a matched system comparison. Add training or infrastructure depth only when your target role requires it.
Start today
- Understand the job: chapters 1–2. Write the user, task, consequence, and tested population.
- Use the tools well: chapters 3–4. Inspect the workspace and write a bounded agent prompt.
- Define and grade: chapters 5–7. Produce a task contract and test the grader's failure paths.
- Run it: chapter 8 and the practice lab. Inspect a complete report and a deliberate invalid-input rejection.
- Review code: chapters 9–10 and 28. Reproduce the coding-lab defect, review a repair, and verify behavior.
- Interpret honestly: chapters 12–14. Write a result with denominator, limits, and a next experiment.
Download the local starter kit
Pick your specialization
| Track | Chapters | What to build |
|---|---|---|
| Health & public health | 15, 18 | Completeness or evidence-state eval with routing checks |
| Medicine | 16 | Faithful extraction with unknown and negation handling |
| Biology & life sciences | 17 | Source-supported evidence extraction or benign analysis validation |
| Biosurveillance | 18 | Temporally faithful backtest with event-level metrics |
| AI safety & biosecurity | 19–20 | Defensive boundary test with benign-usefulness controls |
| Coding-agent evaluation | 9–11, 28–29 | End-to-end repair review and failure packet |
| Research engineering | 21, 27, 29 | Controlled experiment and reliable execution pipeline |
Learn what the current postings ask for
Read chapters 26 and 30. The source review covers selected official roles at OpenAI, Anthropic, the x.ai career site, Google DeepMind, METR, Apollo Research, and Scale. Requirements are distinguished from this course's curriculum recommendations. It is a dated sample, not an exhaustive vacancy inventory.
Download the dated job-source register.
Your readiness checklist
- Can you define what the eval measures and what it does not?
- Can you inspect the code and reproduce a failure?
- Can you defend the labels without seeing model answers?
- Can you identify leakage and a faulty grader?
- Can you run the complete cohort with explicit errors?
- Can you interpret uncertainty and critical failures?
- Can you write a recommendation tied to observed evidence?
The next artifact matters more than another course
When a checklist item fails, return to its lesson and produce the missing artifact. The course includes concepts, exercises, solutions, and templates. External primary documentation remains useful for interfaces that change; no Coursera subscription or university enrollment is needed to use this manual.