Bryan Tegomoh / Learn

Your focused learning syllabus

Learn the common evaluation loop, then specialize. Each stage ends with an artifact you can inspect and explain. Reading progress is a convenience, not a competence score.

TL;DR

Today: understand the system, define a task, run the kit, inspect a failure, and review a coding repair. Next: build a domain reference standard and a matched system comparison. Add training or infrastructure depth only when your target role requires it.

Start today

  1. Understand the job: chapters 1–2. Write the user, task, consequence, and tested population.
  2. Use the tools well: chapters 3–4. Inspect the workspace and write a bounded agent prompt.
  3. Define and grade: chapters 5–7. Produce a task contract and test the grader's failure paths.
  4. Run it: chapter 8 and the practice lab. Inspect a complete report and a deliberate invalid-input rejection.
  5. Review code: chapters 9–10 and 28. Reproduce the coding-lab defect, review a repair, and verify behavior.
  6. Interpret honestly: chapters 12–14. Write a result with denominator, limits, and a next experiment.

Download the local starter kit

Pick your specialization

Track Chapters What to build
Health & public health 15, 18 Completeness or evidence-state eval with routing checks
Medicine 16 Faithful extraction with unknown and negation handling
Biology & life sciences 17 Source-supported evidence extraction or benign analysis validation
Biosurveillance 18 Temporally faithful backtest with event-level metrics
AI safety & biosecurity 19–20 Defensive boundary test with benign-usefulness controls
Coding-agent evaluation 9–11, 28–29 End-to-end repair review and failure packet
Research engineering 21, 27, 29 Controlled experiment and reliable execution pipeline

Learn what the current postings ask for

Read chapters 26 and 30. The source review covers selected official roles at OpenAI, Anthropic, the x.ai career site, Google DeepMind, METR, Apollo Research, and Scale. Requirements are distinguished from this course's curriculum recommendations. It is a dated sample, not an exhaustive vacancy inventory.

Download the dated job-source register.

Your readiness checklist

  • Can you define what the eval measures and what it does not?
  • Can you inspect the code and reproduce a failure?
  • Can you defend the labels without seeing model answers?
  • Can you identify leakage and a faulty grader?
  • Can you run the complete cohort with explicit errors?
  • Can you interpret uncertainty and critical failures?
  • Can you write a recommendation tied to observed evidence?

The next artifact matters more than another course

When a checklist item fails, return to its lesson and produce the missing artifact. The course includes concepts, exercises, solutions, and templates. External primary documentation remains useful for interfaces that change; no Coursera subscription or university enrollment is needed to use this manual.