Daily harness signal

Make fan-out respect the budget it actually has

August 13, 2026 · JST One fresh finding Workflows · CPU · prompt cache
Parallel agents share two scarce resources before they do useful work: local execution capacity and prompt-cache creation. Claude Code 2.1.229 makes both schedulers visible enough to turn into a release gate.
01 · Fresh · source date 2026-08-12

Gate concurrency and cache reuse together

Use when: a dynamic workflow runs in Docker, Kubernetes, a self-hosted runner, or any fan-out where sibling agents receive a long shared rubric. Before 2.1.229, CPU-limited containers could size concurrency from host cores, while simultaneous same-prefix siblings could each pay to create the same prompt cache. The result was misleading parallelism: too many local workers, repeated prefix cost, and rate-limit bursts before useful work began.

Action: make Claude Code 2.1.229 the minimum version for containerized workflows. Declare CPU limits in the runtime instead of relying on node size. Keep the shared instructions byte-identical and at the front of every sibling prompt; append labels, file names, and task-specific text only afterward. Leave CLAUDE_CODE_WORKFLOW_PREFIX_STAGGER_MS unset—the vendor states that zero disables the new staggering. Start with a small /config workflowSizeGuideline=small and raise it only after measuring peak live agents, queue depth, cache creation/read tokens, 429s, and wall time. Do not treat the size guideline as a cap; the current workflow guide calls it advice, while runtime limits remain separate.

Acceptance check: on a high-core host, run the same disposable three-lane workflow in containers capped at four and eight CPUs. Each lane should start with an identical roughly 14k-token prefix and end with a unique sentinel. Record peak live agents and inspect each workflow agent journal under ~/.claude/projects/. Pass only if peak concurrency changes with the container budget rather than staying fixed to host cores, excess lanes queue, all sentinels return, the first same-prefix sibling performs the large cache creation, and later siblings show cache reads instead of equivalent fresh creations. Repeat once with CLAUDE_CODE_WORKFLOW_PREFIX_STAGGER_MS=0; the cache-read advantage should disappear. Keep that negative canary out of production.

Evidence: Anthropic’s official v2.1.229 release explicitly names both repairs: container-limited workflows now use the container CPU limit, and same-prefix siblings are staggered for prompt-cache reuse. Current official workflow docs confirm a maximum of 16 concurrent agents, fewer on CPU-limited machines, queueing beyond the live cap, per-agent token visibility, and advisory—not binding—size guidelines. Issue #63981 supplies an inspectable pre-fix reproduction: three lanes each recreated an approximately 39k-token prefix despite a prewarm call.

Caveat: prefix staggering is a cache optimization, not rate-limit-aware admission control. Open issue #70498 reports heavy-read fan-outs still bursting input-token limits, dropping agents after 429 retries, and forcing manual concurrency reduction. Cache reuse also requires a genuinely shared leading prefix; early per-agent variation defeats it. CPU compliance does not bound nested subprocesses, memory, total agents, or spend, so retain independent budgets and fail-visible completion receipts.

Compact source notes

  1. Claude Code v2.1.229 (published 2026-08-12). Official release note naming the effective-container-CPU fix, same-prefix sibling staggering, and the environment variable that disables the latter.
  2. Claude Code workflow guide (retrieved 2026-08-13). Current authoritative constraints, queueing behavior, token-usage visibility, workflow-size semantics, and runtime-cap distinction.
  3. Claude Code issue #63981 (2026-05-30). Versioned, measured pre-fix reproduction using per-agent cache_creation_input_tokens and cache_read_input_tokens; closed stale without maintainer adjudication.
  4. Claude Code issue #70498 (2026-06-24). Concrete open counterexample showing that agent-count concurrency remains a crude proxy for input-token rate limits.
  5. Docker CPU constraints and Kubernetes resource management (retrieved 2026-08-13). Primary platform documentation confirming that CPU ceilings are enforced through container cgroups rather than host core inventory.
  6. Method: anchor lens—official release semantics, current workflow/runtime docs, a measured pre-fix cache reproduction, and a current rate-limit contradiction; unity lens—effective local capacity and repeated prompt-prefix creation are coupled admission costs in one fan-out. Failure to change peak width with container limits or to convert later prefix creation into cache reads falsifies the gate. Confidence: Likely.