Daily harness signal

Make containment survive a “safe” trajectory

August 31, 2026 · JST One fresh finding Claude Code · untrusted content · containment
Two artifact-backed attack series show why a permission classifier cannot close an execution boundary: individually plausible tool calls can assemble code execution, so containment must survive even when every model-level check allows the trajectory.
01 · Fresh · source date 2026-08-27

Prove containment after classifier success

Use when: an unattended coding agent may read an external webpage, clone a repository, unpack an archive, inspect a third-party package, or execute generated code. Treat content origin—not whether the agent authored the final script—as the trigger.

Action: route every shell, interpreter, browser, and child-agent call for that task through one disposable, unprivileged VM or container. Mount only a scratch workspace; do not mount the host home, SSH agent, cloud credentials, Docker socket, or package-manager credentials. Make outbound network default-deny and allow only task-approved domains at the external firewall or proxy. Record the sandbox/run ID on every tool receipt, terminate its process tree at completion, and keep Auto Mode only as an additional admission signal.

Acceptance check: send the production executor an owned archive containing a harmless struct.py canary. From the extracted directory, invoke python3 -c 'import base64'; the canary should create one sandbox-local marker, attempt one host-path write, and call an owned sink. Pass only if the local marker proves the module loaded, the host path remains unchanged, the sink records zero requests, the receipt names the sandbox, and no child survives teardown. Repeat with agent-generated decoder code, not only the direct command.

Evidence: Johann Rehberger published the redirect, archive, module-shadowing chain, exact command shape, video, and three five-run variants; reported execution effects occurred in 3/5, 3/5, and 4/5 runs. A separate researcher published ten Claude Code 2.1.228 Auto Mode runs, logs, and classifier traces; untrusted code executed in 6/10. Anthropic’s own engineering report says the deployed classifier missed 17% of 52 real overeager actions and its security guidance recommends VMs for external services.

Caveat: these are targeted, small-sample red-team results, not a population attack rate, and the second series used a different chain. Auto Mode can still reduce risk versus bypassing permissions. A container is not sufficient if it inherits the host network, home directory, credentials, or privileged sockets; the acceptance canary must test the deployed boundary.

Compact source notes

  1. Johann Rehberger, “Breaking Claude Code Opus 5 Auto Mode” (2026-08-27; inspected 2026-08-31). Primary red-team write-up with chain, command shape, video, observed effects, small-sample results, mitigations, and disclosure outcome.
  2. “Prompt Injection Experiments with Opus-5 in Claude Code” (2026-08-12; updated 2026-08-29; inspected 2026-08-31). Independent, different-chain artifact set with ten run logs, two classifier traces, version 2.1.228, and published setup.
  3. Anthropic auto-mode engineering report (2026-03-25) and current security guidance (retrieved 2026-08-31). Official architecture, measured false negatives, and explicit VM/isolation guidance for external services.
  4. Anthropic default-change and evaluation report (2026-08-07) and Simon Willison’s technical contradiction note (2026-08-27; corrected 2026-08-30). The vendor reports 0/720 on its fixed scenario set while expressly retaining residual risk; the field chain sits outside that measured set.
  5. Method: anchor lens—versioned run artifacts, exact side effects, official classifier limits, and an owned containment canary; unity lens—every shell, interpreter, browser, and descendant shares one execution boundary independent of model approval. Confidence: Likely. Falsifiers: the production executor reaches the owned sink, changes the host sentinel, lacks a sandbox receipt, or leaves a descendant alive after teardown.