Use solved work as a calibration fixture
Use when
You are choosing which familiar coding tasks to delegate after a model or harness change, and have no local evidence of quality or supervision cost. Prefer a bounded replay over committing to an unattended software factory.
Action
For an explicitly authorized, budgeted pilot, choose already-solved work with a known starting revision. Prepare a disposable environment containing only permitted starting material. Keep the manual solution, later commits, previous transcripts, and answer-bearing artifacts outside every agent-reachable route. A worktree alone is insufficient: shared Git history can reveal the answer. Give the original request, normal authorized tools, and fixed limits; do not smuggle solution hints into follow-up prompts. Freeze the functional, quality, and safety criteria before starting. Record setup, prompting, review, repair, active human time, elapsed time, and agent spend separately.
Acceptance check
Both manual and agent outputs must pass the same independent behavioral checks and review bar; identical diffs are unnecessary. Before launch, inspect reachable files, history, remotes, and memory for solution exposure; afterward, inspect the trace for leakage. Mark contaminated trials invalid, not successful. Preserve failed attempts and all interventions. Keep the task class manual or supervised if quality fails or human overhead exceeds the predeclared acceptable limit. Require held-out similar tasks before expanding delegation; revisit after material model or harness changes.
Evidence
Hashimoto describes recreating manual commits without showing the solution, comparing quality and function, and learning when not to use agents. Anthropic’s evaluation guide independently reports agents gaining an unfair advantage from previous-trial Git history and recommends isolated trials with outcome-based grading.
Caveat
This is a calibration procedure, not a measured speedup or an evaluation run today. Hashimoto’s account provides no controlled paired dataset; the human already knows the solution, so coaching and grading can leak knowledge. Replaying known work teaches task boundaries but does not establish a causal productivity gain. Do not turn his personal learning exercise into a requirement to duplicate every commit or keep agents running without useful work.