Daily harness signal · September 29, 2026

Red for the
right reason.

One implementation lesson · Verification integrity

A deliberately broken build does not prove a regression test works. Require a passing control, the intended failure, and verified restoration before accepting an agent’s mutation-check claim.

Evergreen · source August 9, 2026

Make the negative control diagnostic

Use when

An agent temporarily removes a feature or changes a branch to show that a new test detects the defect. This matters especially for UI tests: broken navigation can prevent the intended assertion from ever running.

Action

Add this contract to the existing verification handoff: “Record revision, environment, mutation, target test/assertion, expected failure, observed failure, and restoration evidence.” In an authorized disposable copy, first pass the unchanged control. Keep tests fixed; change only the targeted production behavior. Capture the actual failure message and file:line, not merely a nonzero exit. Restore only the mutation, compare affected bytes and the complete diff against the pre-mutation snapshot, then rerun affected checks. Preserve unrelated work; never use a blanket reset. A filename-only status check cannot prove restoration.

Acceptance check

An owned pane-visibility fixture must produce: control passes; feature-removal mutant reaches the pane assertion and fails for the expected missing pane; restored code matches the snapshot and passes. Then break earlier navigation: both variants fail early, and the reviewer must reject that as sensitivity evidence. A leftover mutation hunk must block promotion, even if a test passes. Retain raw outputs, revision/environment, and diffs. Independent task acceptance still applies.

Evidence

Superpowers issue 2113 describes correct and mutated builds failing at the same earlier navigation assertion, plus temporary source changes left behind. The tagged 6.4.1 TDD procedure already requires failure for the expected reason; its testing reference names the production break before the test. Stryker’s official documentation separates test failures from runtime and compile errors.

Caveat

The incident is one open, unconfirmed report, without public application code or raw logs. This is our operational adaptation, not a shipped repair. The reference’s mutation check is explicitly mental; do not report it as execution. Line numbers alone are insufficient: inspect the behavioral failure. Timeouts can legitimately detect infinite-loop mutants, so preserve separate outcome categories rather than rejecting all timeouts. No fixture was executed; this newsletter authorizes no evaluation, installation, or skill change.

Compact source notes

  1. seff34’s field report, August 9, 2026, sections 2–3. No retrieved maintainer confirmation or exact deployed version. The restoration paragraph’s green-run wording is ambiguous; no exact execution sequence is inferred from it.
  2. Superpowers tagged TDD procedure, testing reference, and reviewer prompt; release September 19. Retrieved reference history records July 24 changes. These are inspectable instructions and examples, not passing runtime receipts. Blanket deletion and universal full-suite prescriptions are not adopted.
  3. Stryker outcome definitions and configuration, undated, retrieved today: errors are distinct from killed mutants; timeouts count as detected. Default inPlace=false uses a copy; dryRunOnly runs instrumented code without active mutants. This corroborates method, not the reported incident.
  4. Anchor lens: the specific early-assertion failure and inspected source contracts. Unity lens: a negative control must isolate the claimed behavior while preserving the deliverable; proposed invariant. Source contents: confidence_confirmed; incident: confidence_unconfirmed; adaptation: confidence_likely. Accepted wrong-reason failure or residual mutation falsifies the gate.