01. Start here: what this job demands

Published

October 3, 2026

A strong AI evaluator turns an ambiguous claim about a system into an inspectable experiment. The work is to specify what should happen, observe what actually happened, distinguish system failures from evaluation failures, and explain what the evidence supports. A fluent response is an observation, not a verdict.

NoteLearning objectives
  • Define the decision an evaluation must support.
  • Distinguish domain expertise from staff-level engineering experience.
  • Produce the first task, failure, and report artifacts.
TipTL;DR

Your job is to turn a claim into a defensible experiment. Start with one task, one valid grader, one inspected failure, and one report. AI tools accelerate implementation; staff-level judgment requires practice.

This course is designed for a physician-scientist who can bring domain expertise and learn the implementation through AI-assisted practice. It is a working manual, not a promise of staff-engineer qualification. Your clinical and epidemiological judgment transfers especially well to case definition, reference standards, missingness, leakage, uncertainty, subgroup analysis, and harm assessment. Software judgment still needs repeated practice on real systems.

The role in the supplied posting

The supplied G2i posting asks for deep expertise in Python, TypeScript/JavaScript, or Go; staff, principal, architect, or technical-lead experience; hands-on debugging; and detailed assessments of agent behavior. Its explicit requirements make this a different role from medical annotation or judging answers for correctness. The supplied compensation, availability, and onboarding details describe that pasted posting and have not been independently verified as a current vacancy.

Work product What a strong evaluator supplies
A task specification An observable objective, environment, permitted actions, and success criteria
A reproducible run Exact inputs, versions, permissions, timestamps, artifacts, and budgets
A judgment Criterion-specific findings tied to concrete evidence
A failure analysis Root cause, affected cases, plausible alternatives, and uncertainty
A recommendation A decision bounded to the tested system and context

Do not confuse implementing an eval with evaluating an engineering workflow. The former builds the measurement apparatus. The latter may involve studying a model’s interaction with a repository, examining patches, running tests, and deciding whether its approach was appropriate. A role may ask you to do both.

The first-day route

Read chapters 2, 3, 4, 6, 8, 9, and 12. Download the starter kit. Run the scorer and its tests. Complete the coding repair exercise, then write a criterion-based review of the repair. Open the domain chapter closest to your expertise and draft one task contract before building anything else.

The first-day deliverable is small: one reproducible task, one trustworthy grader, one inspected failure, and one report you can explain without an agent. Completing it demonstrates initial competence with the workflow. It does not establish mastery of architecture, security, or staff-level engineering.

Learn by making decisions

For every lesson, ask: what decision would this evidence change? A score with no decision attached becomes a dashboard decoration. A proposed improvement should say what failure it addresses, how the next run will test it, and what result would cause you to reject it.

Use the tools to generate scaffolding, explain code, suggest edge cases, and automate bookkeeping. Retain ownership of the objective, reference standard, access boundary, and conclusion. If you cannot explain why a test passes, ask the agent for an explanation and verify it against a different case.

Your opening exercise

Write three sentences: who uses the system, what action it influences, and what a consequential failure looks like. Then write one sentence naming the population to which your eval result could reasonably apply. If that fourth sentence is broader than your cases, narrow it.

Acceptance criterion: another reviewer can identify the proposed user, task, harm, and tested population without asking you to restate the goal.