17. Biology and life sciences evaluations

Published

October 3, 2026

Biological research evals should test evidence handling and scientific validity rather than confident narration. Start with contained, benign workflows: literature extraction, metadata quality, reproducible analysis, or synthetic trial screening.

NoteLearning objectives
  • Check cohorts, units, and scientific claim support.
  • Inspect analysis reproducibility.
  • Distinguish working software from valid biological conclusions.
TipTL;DR

Test scientific evidence handling, units, cohorts, and reproducibility. Working code can still support an invalid scientific conclusion; keep biological teaching cases benign.

Worked task: evidence extraction

Supply a short synthetic study description with a stated cohort, endpoint, comparator, and limitation. Ask the system to extract those fields and support each with a text span. Include a case in which the paper reports an association but not a causal experiment. The expected output must preserve that distinction.

Evaluate correct entities, study design, denominators, endpoint status, and uncertainty. Add negative controls: an absent endpoint, a mismatched comparator, a secondary outcome presented as primary, and a preprint mistaken for an independently replicated clinical result.

Trial eligibility exercise

Use a synthetic protocol, not real trial guidance: criterion A must be explicitly true; criterion B must be explicitly false; unknown status requires review. The output is one of eligible, ineligible, or needs_review, with criterion-level support.

The policy logic should be visible. Never infer an exclusion from missing data. When a patient record and protocol have incompatible dates or definitions, the output should surface the discrepancy. Real eligibility requires the current protocol, institution-specific process, and qualified review.

Scientific software evaluation

Evaluate whether a coding agent preserves units, sample IDs, cohort membership, missing-data policy, and analysis intent. A program that runs can still produce a scientifically invalid result. Tests should include known analytical fixtures and independent numerical checks.

For benign genomics metadata, test sample joins, duplicate identifiers, reference-version handling, and missing fields. Do not infer epidemiological linkage from a matching identifier alone. Distinguish metadata validation from biological or clinical interpretation.

Reproducibility

Record inputs, preprocessing, excluded records, package versions, random seeds, and intermediate outputs. Separate software reproducibility from validity of the scientific conclusion. Reproducing a biased analysis repeats the bias.

For a research agent, score source retrieval, accurate extraction, citation-to-claim fit, and acknowledgment of conflicting evidence separately. A DOI that resolves is not enough: the linked paper must support the claim and population.

Safe boundaries

Use non-actionable, benign research cases for public teaching. Dangerous biological capability evaluations require institutional authorization, specialized review, controlled access, and restricted artifacts. Public examples should measure defensive behavior without providing protocols that increase harmful capability.

Exercise: ask a coding agent to calculate a descriptive rate from synthetic counts. Give one record with a zero denominator and one with incompatible units. Judge whether it fails loudly, flags the inconsistency, or silently produces a number.