Make grader approval evidence-backed
Use when
An agent builds an evaluator or requests label approval. Hamel Husain’s leasing-assistant walkthrough found premature failure selection, aggregate counts without enough context, and four different checks bundled into one transfer verdict.
Action
Add this gate to the existing workflow: “First review actual failures. For each proposed graded case, show its ID, input, relevant full trace, grader code or prompt, individual check results, and proposed label. Record corrections against that ID and grader revision. Do not start optimization until review is complete.” Separate deterministic event-order checks from semantic judgments such as whether consent was given. Reuse an existing safe viewer or readable records; a new web app is not required. Without a reviewer, leave a review-pending artifact, not self-approved grades.
Acceptance check
In an authorized offline fixture, provide consent-then-transfer and transfer-before-consent traces with deliberately swapped proposed labels. The reviewer must see both sequences and correct the right cases; unresolved mislabels cannot enter the approved set. Remove one trace: approval must remain blocked. Reopen the records: corrections, grader revision, and pending/approved status must persist. A correct positive control must remain approvable. Retain the review receipt; summary counts alone fail this gate.
Evidence
Husain documents specific questions and grader conditions, not just a product opinion. Anthropic’s September 28 announcement describes scored-trace review; September 29’s merged source already requires literal signoffs and full-exchange or trace-link inspection. That contrast matters: test the delivered review experience, rather than assuming written instructions were followed.
Caveat
This is one practitioner session; its installed revision is unspecified. A promised update is not a shipped repair. Inspectable records do not establish grader accuracy: calibrate on fresh, independently human-labeled cases before trusting optimization scores. Preserve independent task-outcome acceptance, privacy, safe rendering, and separate execution-budget approval. Our fixture was not run. This newsletter authorizes no evaluation, installation, or skill change.