Bryan Tegomoh / Learn
Chapter 23

Calibration interview and readiness assessment

TL;DR

Practice explaining the system, reference, grader, failures, and limits. Staff-level coding evaluation needs demonstrated engineering judgment beyond reading this course.

The supplied posting describes a technical review intended to mirror actual work. Prepare by doing that work and explaining the evidence. Do not represent AI-generated code as personal experience you do not have.

Questions you should answer clearly

  1. What exactly is the system under test?
  2. How did you choose cases and prevent leakage?
  3. Why is the grader valid for this construct?
  4. What does a passing test fail to establish?
  5. How do you distinguish a broken task from a weak agent?
  6. What would cause a correct-looking patch to fail in production?
  7. Which failure matters most and why?
  8. How would you test whether a judge prefers verbose answers?
  9. What changes when the model has tools and memory?
  10. What would you refuse to conclude from your results?

Model answers to practice

A model has 95% accuracy. Can we deploy it? Not from that statement. Define the task, sampling, denominator, comparator, critical failures, uncertainty, workflow, and deployment conditions. A single percentage does not answer those questions.

The patch passes tests. Is it correct? It satisfies the tested behavior under that environment. Inspect whether the tests cover the task contract, preserve existing behavior, and exclude implementation-specific assumptions. Examine untested failure paths and side effects.

Can AI tools replace your engineering knowledge? They can accelerate writing and investigation. I still need to inspect interfaces, tests, state, permissions, and failure mechanisms to judge whether the work is valid. My domain expertise is strongest where the task requires clinical or epidemiological reference standards.

What if the reference answer is wrong? Preserve the case, document the discrepancy, obtain adjudication, version the correction, and recompute affected scores. Do not silently change it after seeing which system benefits.

Practical assessment

Give yourself an unfamiliar small repository and a task contract. Before using an agent, write expected behaviors and two likely failure paths. Then observe the agent, inspect its patch, execute checks, and write a review. Compare your findings with a qualified engineer's review where available.

Readiness levels

You are ready to assist with supervised eval work when you can execute the pipeline, follow security requirements, and write concrete findings. You are ready to own a domain eval when you can defend its reference standard, design, and limitations. Staff-level coding evaluation additionally requires demonstrated engineering judgment across complex real systems.

The goal is evidence of competence, not a self-awarded title. Where the G2i role is beyond your current engineering experience, domain evaluator, evaluation scientist, clinical AI evaluation, or supervised eval engineering can be a more credible entry point while you build the coding-review portfolio.