Daily harness signal

Make the bug
replayable.

September 9, 2026 · JST · One evergreen implementation lesson
When an agent cannot reliably see a dynamic bug, change the feedback surface before changing the prompt: one replayable case, shared application logic, fast text diagnostics, and a visual acceptance check.
Evergreen · source date February 5, 2026

Give humans and agents the same experiment

Use when: animation, UI state, or sequence-dependent behavior makes “fix this bug” hard to verify. Lewis Metcalf’s Amp example starts with a ball teleporting through a rotating ring: randomly timed screenshots hide the causal sequence.

Action: extract the existing state-transition logic into one module used by both the application and a small replay harness; do not write a second simulation. Encode the reproducer’s inputs in a URL or fixture, including initial state, seed, step count, and timestep where relevant. Render a static trajectory for inspection, then expose the same inputs through a headless command that prints bounded frame logs. The article’s example is node physics-cli.js --vx=-7.71 --vy=2.14 --from-frame=17 --to-frame=25 --delta. This is a project-local interface example, not an installed utility. Prompt: “Reproduce this fixture before editing. Use frame deltas to isolate the failure; after repair, replay the same case and neighboring inputs, then show the visual result.”

Acceptance check: freeze the bad fixture and expected behavior before repair. Repeated runs must reproduce the same trace on the same build. Require agreement between headless state and visual playground at selected frames. The original build must fail the agreed oracle; the repaired build must pass it without changing the fixture or weakening expectations. Nearby edge cases and the actual application must also pass. Preserve before/after logs, images, inputs, and source revision.

Evidence: Metcalf’s dated engineering post publishes exact velocity URLs, a shared-module CLI sketch, commands, and frame-level output. Its delta trace makes an edge-collision discontinuity inspectable rather than merely described. OpenAI’s separate harness report supports exposing UI, logs, and metrics directly to agents; it does not independently validate this physics repair.

Caveat: text diagnostics accelerate localization, not proof. A clean CLI trace can miss rendering, timing, integration, or nondeterministic failures. The published CLI sketch omits argument parsing and logging details; it is not a complete drop-in package. No demo, code, or runtime test was executed for this newsletter. Keep the harness small and retain human visual acceptance.

Compact source notes

  1. Lewis Metcalf / Amp, “Feedback Loopable” — February 5, 2026. Primary workflow, parameterized cases, CLI sketch, and diagnostic output inspected, including the complete delta-log section. The author describes a fixed interactive demo; its behavior was not tested here.
  2. Ryan Lopopolo / OpenAI, “Harness engineering” — February 11, 2026. Independent engineering account of agent-visible application and observability surfaces, not controlled evidence of speedup from this particular recipe.
  3. Thorsten Ball, “Joy & Curiosity #73” — February 8, 2026. Credits colleague Lewis; the legacy queue ID is preserved for deduplication, with authorship corrected. This recommendation is a discovery lead, not independent validation.
  4. Anchor lens: exact reproducer inputs and inspectable frame logs. Unity lens: one transition implementation, multiple observation surfaces. Proposed acceptance control: confidence_likely. Falsifier: headless acceptance passes while the same case remains broken in the actual application.