Daily harness signal · September 16, 2026

Replay before
you delegate.

One implementation lesson · Calibration without answer leakage

A model upgrade is not evidence that your work became delegable. Replay familiar tasks without exposing the solution, then judge accepted outcomes and human supervision—not how impressive the session feels.

Evergreen · source February 5, 2026

Use solved work as a calibration fixture

Use when

You are choosing which familiar coding tasks to delegate after a model or harness change, and have no local evidence of quality or supervision cost. Prefer a bounded replay over committing to an unattended software factory.

Action

For an explicitly authorized, budgeted pilot, choose already-solved work with a known starting revision. Prepare a disposable environment containing only permitted starting material. Keep the manual solution, later commits, previous transcripts, and answer-bearing artifacts outside every agent-reachable route. A worktree alone is insufficient: shared Git history can reveal the answer. Give the original request, normal authorized tools, and fixed limits; do not smuggle solution hints into follow-up prompts. Freeze the functional, quality, and safety criteria before starting. Record setup, prompting, review, repair, active human time, elapsed time, and agent spend separately.

Acceptance check

Both manual and agent outputs must pass the same independent behavioral checks and review bar; identical diffs are unnecessary. Before launch, inspect reachable files, history, remotes, and memory for solution exposure; afterward, inspect the trace for leakage. Mark contaminated trials invalid, not successful. Preserve failed attempts and all interventions. Keep the task class manual or supervised if quality fails or human overhead exceeds the predeclared acceptable limit. Require held-out similar tasks before expanding delegation; revisit after material model or harness changes.

Evidence

Hashimoto describes recreating manual commits without showing the solution, comparing quality and function, and learning when not to use agents. Anthropic’s evaluation guide independently reports agents gaining an unfair advantage from previous-trial Git history and recommends isolated trials with outcome-based grading.

Caveat

This is a calibration procedure, not a measured speedup or an evaluation run today. Hashimoto’s account provides no controlled paired dataset; the human already knows the solution, so coaching and grading can leak knowledge. Replaying known work teaches task boundaries but does not establish a causal productivity gain. Do not turn his personal learning exercise into a requirement to duplicate every commit or keep agents running without useful work.

Compact source notes

  1. Mitchell Hashimoto: My AI Adoption Journey — February 5, 2026, especially “Reproduce Your Own Work.” First-person method and documented limitations, not a controlled benchmark. The bounded pilot, leakage audit, accounting fields and promotion criteria above are our adaptation.
  2. Anthropic: Demystifying evals for AI agents — January 9, 2026; “Build a robust eval harness” documents Git-history leakage, while “Design graders thoughtfully” warns against rejecting valid alternative paths. Method-level corroboration, not replication of Hashimoto’s exercise.
  3. Anchor lens: inspected replay procedure and specific contamination failure. Unity lens: calibrate delegation against unchanged outcomes with complete human-cost accounting; proposed invariant, not demonstrated cross-harness equivalence. confidence_confirmed for source contents; confidence_likely for this untested adaptation. Accepting a leaked answer, failed outcome or concealed repair effort falsifies the proposed gate.