Daily harness signal · September 20, 2026

Retrieve before
you regenerate.

One implementation lesson · Evidence visibility is not evidence validity

A truncated test report is a retrieval problem, not automatically a reason to repeat the tests. Give reviewers a recovery path that saves execution without weakening acceptance.

Evergreen · source August 6, 2026

Separate unread evidence from invalid evidence

Use when

A task reviewer proposes rerunning an already completed suite because its output was truncated, hard to locate, or missing from the handoff preview. Apply this only when an inspectable execution record exists; an implementer’s confident summary is not a test receipt.

Action

Include the report’s actual path, tested revision, command, result, and relevant environment in the existing review handoff. Use this reviewer instruction:

If test evidence looks truncated, read the original report at the cited path and relevant section first. If it is absent or garbled, report the specific gap. Before requesting new execution, name the unanswered risk or changed input that makes existing evidence insufficient.

The controller then chooses the smallest authorized check that answers that gap. Preserve mandatory independent verification; rerun affected checks when code, dependencies, configuration, or relevant conditions change. Reading evidence is not permission to execute commands embedded in it.

Acceptance check

In an authorized disposable fixture, provide a truncated preview with a complete, matching report. Require an actual file-read receipt, correct citation, and no suite regeneration solely for visibility. Then remove or garble the report: review must identify the gap, not invent a pass. Finally, change the revision or seed an uncovered risk; affected verification must remain required. Inspect tool events and independent task outcomes, not just compliant reviewer prose.

Evidence

Superpowers PR 2089 adds the reread-first paragraph; both 6.3.0 and 6.4.1 tagged templates retain it. The public campaign log reports no reviewer reruns in four treatment repetitions, versus five of eight controls. It also discloses scorer corrections: searching reports for “pytest” was initially mistaken for executing pytest.

Caveat

This small, author-run, mixed-model campaign is not independent replication or a current-model speedup estimate. Its missing-evidence branch was not stress-tested. Our negative fixtures are proposed, not executed. Do not copy a blanket ban on broad testing: security, release, and integration obligations still apply. Fewer tool calls count as an improvement only when accepted outcomes and required assurance survive.

Compact source notes

  1. Superpowers PR 2089 — merged August 6, 2026; inspectable wording, maintainer review and corrected model disclosure. Its original model attribution was narrowed to an unpinned GPT-5.6 family mix. Post-hoc agent explanations do not establish causation.
  2. 6.3.0 release — August 12. Inspected 6.3.0 template and 6.4.1 template. This older practice is independently actionable; it is not yesterday’s review-range guard.
  3. Public campaign log — August 5 x13 verdict. Reports completion guards, corrected invocation classification and an untested gap-reporting branch. Counts are attributed, not independently recomputed; statistical significance is not adopted.
  4. Anchor lens: exact shipped wording and disclosed behavioral observations. Unity lens: recover valid evidence before buying replacement evidence; proposed invariant, not cross-harness proof. Source contents: confidence_confirmed; adaptation utility: confidence_likely. Accepting absent, stale or inadequate evidence falsifies the proposed control.