Use Codex, Claude Code, and Cursor effectively
TL;DR
Give coding agents a bounded contract and demand executed evidence. Use them for implementation and investigation while you own the objective, reference, permissions, and conclusion.
Give a coding agent an operational contract, not a motivational request. The useful unit of delegation is a bounded task with evidence of completion. Large vague prompts encourage large vague changes.
A prompt you can use today
Read the repository instructions, README, dependency configuration,
relevant implementation, and tests before editing.
Goal: implement a scorer for the supplied task specification.
Constraints: Python 3.12+, strict types, no model calls, no new runtime
packages, no real patient data, no changes outside the scorer and tests.
First explain the input/output contract and likely failure cases.
Implement the smallest complete solution. Add meaningful positive and
negative tests. Run tests and type checking. Show the exact diff.
Report commands executed, observed results, and remaining limitations.
Do not weaken tests or silently skip missing records to obtain a pass.
Change the permitted files and constraints to match the repository. Avoid telling the agent to finish “at any cost,” to bypass approvals, or to make everything green. That instruction invites measurement shortcuts.
A productive interaction loop
First ask for a repository map. Next ask for a plan tied to files and failure paths. Then authorize the bounded implementation. Inspect the diff, run the checks yourself, and ask for a review that focuses on concrete defects. A separate model or session can provide a different perspective, but is not automatically independent evidence.
For diagnosis, ask for multiple hypotheses with discriminating tests. “The API is broken” is a conclusion. “The response parser rejects null while the API permits null; here is a failing fixture” is an inspectable hypothesis.
Tool-specific habits
In Codex, use the workspace context, review the proposed changes, and keep the task scoped to the intended directory. In Claude Code, use repository instructions and the transcript of executed commands to check what it did. In Cursor, inspect changes through its review interface and also verify the repository diff. Interface labels and permissions evolve; confirm current behavior in official documentation before relying on an automation setting.
Official entry points: Codex learning resources, Claude Code engineering guidance, and Cursor documentation. These links are documentation, not endorsements of a particular model configuration.
Ask for evidence, not reassurance
Useful follow-ups include: “Show a case the previous implementation fails”; “Which behavior is not covered by these tests?”; “What input could make this return a misleading success?”; “Which assumption is encoded in this threshold?”; and “What would distinguish a model failure from a broken grader?”
Do not ask an agent to grade its own work and then treat the grade as acceptance. Make it supply artifacts: a test fixture, command output, exact diff, and concrete reasoning that you can inspect.
When to intervene
Stop the run if it expands scope, downloads unexplained executables, accesses private data unnecessarily, changes a reference answer, disables a test, or claims a result it has not executed. Pause for missing credentials or unapproved spending rather than asking the agent to invent a workaround.
Exercise: take a broad prompt such as “build a medical benchmark.” Rewrite it into a single extraction task with specified fields, synthetic inputs, one scoring rule, and a failure test. The smaller prompt should create a more meaningful artifact.