Daily harness signal · September 18, 2026

Make helpers
earn their place.

One implementation lesson · Reversible retirement, not blind cleanup

An optional workflow helper can outlive the problem it solved. Before adding another command, check whether existing automation still improves accepted work—and retire only what you can safely replace.

Evergreen · source July 30, 2025

Give optional automation an exit criterion

Use when

A slash command or prompt wrapper is repeatedly bypassed, endlessly tuned, or adds no demonstrated value beyond a direct request. Exclude security controls, required verification, and rare but important recovery procedures; low usage alone is not a retirement case.

Action

In the existing project record, write the helper’s trigger, actual use, promised benefit, dependencies, and rollback. For an explicitly authorized, budgeted pilot, compare fresh sessions with and without that helper, keeping tasks, context, model, tools, and safety policy fixed. For a bug-fixing wrapper, the alternative is the same issue and relevant context in a direct request—not less information. Predeclare the benefit required to justify retention: better accepted outcomes or lower human effort without quality loss. Record repair work, elapsed time, and usage. Vary only the optional helper; do not disable enforcement to make the baseline faster.

Acceptance check

Both conditions must pass the same independent task checks and required negative safety tests. Inspect traces for accidental helper loading; retain failures, not just successful examples. Keep the helper if it earns its predeclared benefit. If it does not, and the alternative preserves outcomes, propose reversible deactivation with source history intact. After approval, verify ordinary tasks, dependent workflows, and rollback. Inconclusive results remain inconclusive. Never delete live dependencies or mandatory controls as “unused.”

Evidence

Ronacher names failed commands, explains why his bug-fixing wrapper added little, and publishes an acceptability-based repeat-run procedure. His later Pi account still discards unneeded skills while retaining useful review, todo, and interception extensions. The lesson is selective retention, not “automation never works.” Anthropic’s evaluation guide supports repeated trials and grading actual outcomes.

Caveat

This is an untested retirement procedure, not a measured speedup or cleanup performed today. Ronacher’s three repetitions are a smoke test, not statistical confidence. His historical hook limitations are outdated: current Claude Code documents blocking and richer lifecycle events. Do not copy the old permission-bypass or backup-deletion examples. No skill, hook, or configuration was changed.

Compact source notes

  1. Armin Ronacher: Agentic Coding Things That Didn’t Work — July 30, 2025. Named failures, exact command examples and repeated “would I accept it?” inspection; first-person evidence, not a controlled dataset.
  2. Ronacher: Pi, The Minimal Agent Within OpenClaw — January 31, 2026. Later useful extensions and continued selective disposal contradict a blanket rejection of automation.
  3. Anthropic: Demystifying evals for AI agents — January 9, 2026. Trial variability and environment-outcome definitions; method support, not replication. Current hooks reference — undated, inspected today; PreToolUse blocking and lifecycle events supersede the older limitations.
  4. Anchor lens: inspectable workflow examples and dated failure accounts. Unity lens: optional complexity must justify itself against unchanged outcomes; proposed invariant, not measured cross-harness equivalence. Source contents: confidence_confirmed; adaptation utility: confidence_likely. A retired helper’s required behavior disappearing, or a safety gate being weakened, falsifies the proposed acceptance.