20. Privacy, security, and operational discipline
The supplied role describes managed access, security tooling, and NDA workflows. Treat those requirements as part of doing the job. An evaluator who produces accurate feedback while mishandling confidential artifacts is not doing acceptable work.
- Classify sensitive data and enforce least privilege.
- Protect secrets, logs, and hidden references.
- Establish authorized run limits.
Classify data, enforce least privilege, protect logs, and establish budgets before paid runs. Use synthetic cases for learning and institutionally authorized arrangements for real sensitive data.
A data classification plan
Identify public materials, internal task content, restricted customer code, patient data, secrets, and sensitive safety cases. Specify storage, access, retention, approved processing tools, and publication permissions for each. Keep the most restrictive applicable rule when a record combines categories.
Do not paste customer repositories or patient records into a personal AI account without authorization and an appropriate institutional arrangement. Removing names alone is not a reliable de-identification method. HHS guidance describes HIPAA de-identification approaches; institutional privacy and legal review must determine the applicable requirements for real data.
Least privilege
Give a test agent only the files, tools, network destinations, and credentials needed for its task. Keep hidden references outside the agent’s readable environment. A folder name such as “private” is not access control. Verify denied access rather than assuming isolation.
A container helps package and isolate software, but it is not automatically a complete security boundary. Review mounts, network access, privileges, and resource limits. Never mount your personal home directory into an untrusted agent experiment.
Secrets and logs
Use approved credential management. Redact secrets at logging boundaries and test the redaction with synthetic examples. Logs can contain patient text, access tokens, source code, or dangerous content even when the final report does not. Apply the same data classification to traces and screenshots.
Cost and cancellation
Set case limits, per-case token or action budgets, timeouts, concurrency, and a total spend ceiling before a paid run. Keep cancellation and partial-run status explicit. A budget cap that only warns after spending is not enforcement.
Retries and judges add cost. Estimate from expected input/output tokens and current provider prices, clearly labeled as an estimate. Record observed billing separately. Never hardcode a price from memory into a decision report.
Onboarding checklist
Confirm the employer’s approved device, identity provider, access channels, permitted model providers, repository rights, evidence-storage location, and reporting process. Ask how to handle a security incident or an unexpectedly sensitive case. Do not disable security software to make an eval run faster.
Exercise: inspect the starter kit archive and explain why it is safe to publish: synthetic cases, no credentials, no patient data, and no proprietary traces. Then identify what would have to change before publishing a real employer evaluation packet.