Building and Evaluating AI Systems

Practical engineering and evaluation for health, science, and safety

An applied handbook with runnable evaluation exercises, engineering foundations, and specialization in medicine, health, biology, biosurveillance, and AI safety.
Author

Bryan Tegomoh, MD, MPH

Published

October 3, 2026

A practical handbook for building AI systems, measuring their behavior, and interpreting the evidence. The examples connect engineering practice with medicine, health, biology, biosurveillance, and defensive AI safety. Coding agents accelerate implementation; understanding the system and verifying its behavior remain necessary.

NoteLearning objectives
  • Build and inspect a complete small evaluation.
  • Diagnose whether a failure comes from data, model behavior, tools, or measurement.
  • Produce a reproducible artifact and a recommendation with explicit limits.
TipTL;DR

Start with the syllabus, run the starter kit, and inspect a failure. Learn the foundations through executable examples. Progress when an artifact can be produced, tested, explained, and bounded, not when a chapter is merely read.

Choose a learning route

Start with evaluation

Define a task, run the kit, and review a coding repair.

Follow the focused syllabus

Build the foundations

Inspect Python, train a small classifier, test retrieval, and enforce tool permissions.

Begin the runnable foundations

Specialize in health and science

Apply references, temporal controls, uncertainty, and failure severity.

Health and scientific evaluation

Prepare for a technical role

Translate selected official postings into portfolio evidence.

What employers ask for

Learn by running the work

Download the complete starter kit. It contains a strict evaluation runner, authored synthetic cases, tests, task and report templates, a deliberately defective coding exercise, and runnable foundations. It uses Python’s standard library and makes no model API calls.

Open the Practice Lab for commands, a confusion-matrix calculator, and review exercises. Read the complete printable manual or download its Markdown.

What competence looks like

A defensible task contract, validated references, a tested grader, a complete run, an inspected failure, and an evidence-based recommendation. A small reproducible artifact is stronger evidence than a certificate or an unverified demo. Staff-level work additionally requires sustained architecture, debugging, reliability, and collaboration experience.

Scope and sources

The hiring review is a selected sample retrieved on October 2, 2026, not a complete vacancy inventory. The source register preserves the reviewed roles. The foundations incorporate the useful training sequence from an earlier Clinical AI MTS curriculum, with runnable exercises replacing unfinished examples. Model calls and real clinical validation require additional configuration, authorization, and evidence beyond these local teaching exercises.