40. Day-to-day work in AI engineering and evaluations

Published

October 3, 2026

Practical work is organized around evidence-producing tasks: an adapter that ran, a reference that can be defended, a failure that can be reproduced, or a comparison that changes a decision. A working day often alternates between reading traces, changing a small component, running tests, and discussing what the result means. The schedules below are illustrative work patterns, not verified descriptions of a particular company’s employees.

NoteLearning objectives
  • Map technical roles to concrete daily tasks and deliverables.
  • Use coding agents under a bounded implementation contract.
  • Produce an experiment packet whose conclusions can be independently checked.
TipTL;DR

Own an inspectable artifact: a validated grader, working adapter, reproduced failure, or controlled comparison. The worked days show how technical roles turn traces and experiments into decisions, with bounded agent prompts and a complete evidence packet.

Introduction

Use lesson 26 for the dated hiring sample. Use this lesson to translate a role into concrete work. The important question is what artifact the person owns and how another person can verify it.

What each role owns

Role Typical working task Inspectable deliverable
Evaluation engineer Implement a task, grader, or runner; investigate score drift Validated cases, complete run, failure packet
AI application engineer Connect models, retrieval, tools, and user workflow Working feature, tests, trace, measured task outcome
Research engineer Run and debug experiments; improve throughput and reproducibility Controlled experiment, preserved outputs, comparison
Research scientist Form hypotheses, define measurements, interpret experiments Falsifiable question, result, limits, next experiment
Safety or security researcher Probe threats and test controls Threat model, contained reproduction, mitigation evidence
Domain evaluator Define references and adjudicate clinical or scientific outputs Rubric, evidence spans, disagreements, severity findings
Staff engineer Resolve architectural bottlenecks and coordinate system ownership Decision record, migration/recovery plan, verified system behavior

Teams distribute these responsibilities differently. A title does not establish the exact scope, and domain expertise does not remove the need to learn engineering failure modes.

Worked day: an evaluation score drops

Incoming issue: yesterday’s extraction run had fewer supported answers than the preceding run. The provider, prompt, retriever, parser, and reference dataset could each explain the difference.

Start by reproducing one case with the exact saved input and configuration. Compare the raw response with the normalized output. If the parser discarded a valid answer, the observed decline is not necessarily a model regression. If the source set changed, a matched model comparison has not yet been performed.

Ask the coding agent to identify the first divergent boundary:

Inspect the saved baseline and candidate run manifests and one failed case.
Trace input -> eligible context -> raw response -> parsed artifact -> grade.
Do not edit references or acceptance criteria.
Report the earliest material difference with file and case evidence.
Implement the smallest repair only after the mechanism is reproduced.
Run the relevant failure tests and the complete affected cohort.
Return the diff, exact commands, exit statuses, and remaining uncertainty.

An effective end-of-day note is narrow: “The parser rejected a permitted null value. The repair passes absent-value and invalid-type cases. The complete cohort was rerun under the same model configuration. Model behavior has not been shown to improve.” State the actual observed evidence in a real report; this sentence is an example, not a result from an employer run.

Worked day: a safety researcher examines an agent

The task is to assess whether a synthetic retrieved document can induce an unauthorized write. Establish a clean baseline that completes the read task. Insert a harmless redirection while keeping the source’s task-relevant information unchanged. Verify what text reached the model; an injection outside retrieved context is not a tested exposure.

Run under the approved boundary. Save the proposed calls, executed calls, denied calls, final output, and mock destination state. Review benign controls alongside attack cases. Separate the result “the runtime prevented writing” from “the model did not attempt writing.”

The deliverable is a case packet, not a spectacular transcript. Include task contract, threat model, exact fixture, configuration, rubric, and mitigation comparison. Keep sensitive real-world failures within the authorized reporting process.

Worked day: a research engineer optimizes an experiment

Reproduce the baseline before changing its code. Identify the bottleneck with observed timings or profiling. Compare one change under fixed data and hardware, and check correctness independently. Preserve failed runs and intervention records.

An agent can implement candidate optimizations or create plotting scripts. The engineer still verifies whether a faster result skipped work, changed precision, leaked acceptance labels, or altered the computation. The gate in lesson 39 shows why a better score alone is inadequate.

The end artifact is an experiment card with a reproducible command and an appropriately bounded comparison. If the result is inconclusive, record the evidence needed next. Do not turn uncertainty into a performance claim.

Worked day: a biosurveillance benchmark designer

Begin with the decision-time snapshot and the event ledger. Review whether syndicated reports represent one event, whether a correction was available at the tested cutoff, and whether the reference distinguishes unverified reports from actionable evidence.

Implement a deterministic baseline before adding an LLM. Run the temporal replay in lesson 41. Inspect a missed event, an unsupported escalation, and a duplicate queue item. Discuss the consequence and reviewer workload with the operational owner.

The domain specialist’s distinctive contribution is often the reference policy: what counts as the same event, what evidence merits review, which time defines readiness, and what the system must leave unknown. The engineer makes those decisions observable in data structures and tests.

The smallest complete experiment packet

Use the existing templates in lesson 24. One canonical packet contains:

task-contract: intended use, population, consequence, permitted actions
cases: model-visible inputs with provenance and version
references: independent labels, spans, adjudication, event linkage
run-manifest: system, harness, tools, permissions, budget, exact revision
raw-output: complete responses and tool traces
graded-results: per-case outcomes, errors, critical findings
failure-notes: reproduction, mechanism, representative evidence
decision: supported action, limits, unresolved gaps

In a production environment, references and sensitive outputs belong in restricted storage, not necessarily in the same folder. The packet names its artifacts; access policy determines where they reside.

A reliable agent-assisted working routine

Inspect the repository and instructions. Record the baseline and assign file ownership. Ask the agent to implement a bounded piece of work with success and failure checks. Inspect the diff and run the checks. Read representative artifacts yourself. Ask another review process a specific question when independence matters.

Split “implement a validator” from “decide what the reference should mean.” The former can be delegated under a precise contract. The latter needs domain and measurement judgment. An agent proposing a rubric does not make that rubric valid.

Progress should be visible in artifacts. A long session transcript is not a deliverable unless it contains the necessary evidence and can be navigated. Summarize each completed task with the tested revision, affected cases, failure behavior, and next decision.

When work stops for a real reason

A failed provider call is an execution issue. A contradictory reference is a measurement issue. Lack of authorized data is an access issue. A disputed clinical acceptance threshold is a governance decision. Name the category and preserve useful completed work rather than returning a half-run as success.

Do not fill every gap with another framework. If a typed function and a test solve the problem, retain that small design. Add queues, containers, multi-agent coordination, or distributed training when observed requirements justify their maintenance cost.

Practice assignment and self-assessment

Complete one packet using the surveillance or security exercises. Change one fixture deliberately, reproduce the failure, and write the result without overstating what the local program measures. Then use a coding agent to implement one bounded repair under the earlier contract.

A reviewer should be able to answer: what was tested, what changed, what failed, what ran, and what conclusion is justified? If those answers require trusting the agent’s final message, the packet is incomplete.

Checkpoint: select a role from the table, produce its smallest deliverable, and explain the difference between implementation success and evidence that the intended workflow improved.