Daily harness signal

Repair the first broken transition

August 18, 2026 · JST One evergreen finding Evals · traces · localization
End-to-end success tells you whether an agent failed; the first failed transition tells you where to repair it. Keep those layers separate, then require both to improve.
01 · Evergreen · source date 2025-09-23

Turn failed traces into one repair queue

Use when: a multi-step agent has enough complete traces to show recurring failures, but dashboards of tool accuracy, latency, or handoff quality do not identify which harness change should come first. Define the user-level outcome before examining steps: the answer is correct, the side effect exists, or the artifact passes its independent oracle. Only failed runs enter localization.

Action: For every run, persist run_id, terminal outcome, ordered states, tool receipts, and recovery events. Assign exactly one first-failure pair: (last_successful_state, attempted_state). Aggregate those pairs into a transition matrix; then rank cells by frequency times impact, not frequency alone. Select the highest-priority cell and repair only its governing layer—prompt, tool schema, router, handoff, policy, or feedback. Freeze a held-out trace set before the change and rerun the same tasks and budgets afterward. Preserve full traces as evidence, but do not turn the observed path into a mandatory sequence: a different route that reaches the verified outcome remains valid.

Acceptance check: On the held-out fixture, pass only if every end-to-end failure contributes exactly one matrix cell, every successful run contributes none, and the matrix total equals the failed-run count. After the repair, require both the selected cell’s count and the overall outcome-failure count to decrease under identical task inputs, model version, tool surface, and budget; require all safety and side-effect oracles to remain green. Include at least one valid alternate-path fixture and verify it passes despite using different intermediate states. If labels disagree, adjudicate against the raw trace and record the decision before comparing runs.

Evidence: Hamel Husain’s practitioner guide specifies the two-phase order—black-box task success first, then step diagnostics—and gives a concrete transition matrix whose rows are last successful states and columns are first failures. Anthropic independently distinguishes outcomes from transcripts, recommends outcome graders before trace analysis, and warns that fixed tool-call sequences are brittle. OpenAI’s current agent-eval docs likewise use traces to localize workflow failures and repeatable datasets to test changes.

Caveat: First-failure labeling is an analytical judgment, not causal proof. Cyclic, parallel, or self-recovering workflows may need milestone-specific matrices or multiple hypotheses, and rare catastrophic failures can outrank the busiest cell. This method is weakened if independent reviewers cannot reproduce the selected cell, or if its count falls while end-to-end success, safety, or cost worsens.

Compact source notes

  1. Hamel Husain, “How do I evaluate agentic workflows?” (2025-09-23). Practitioner method with an explicit two-phase order and transition-failure matrix; the source links Bryan Bischof’s sequential text-to-SQL example.
  2. Anthropic, “Demystifying evals for AI agents” (retrieved 2026-08-18). Official definitions of transcript versus outcome, grader selection guidance, transcript-review procedure, and warning against brittle fixed-path grading.
  3. OpenAI, “Evaluate agent workflows” and “Trace grading” (retrieved 2026-08-18). Current official workflow for localizing behavior in traces, then promoting findings into repeatable datasets and eval runs.
  4. Method: anchor lens—one outcome oracle, complete run traces, first-failure labels, and held-out before/after counts; unity lens—outcome evaluation and transition localization are separate layers joined by stable run identity. A non-reproducible hotspot or worsening endpoint result falsifies the proposed repair. Confidence: Likely.