Assessing OpenAI
Two fresh fail-closed transition gates for unattended permission prompts and uncertain reconnect state.
Two fresh release gates for policy-bound approval review and bounded MCP process lifetimes in agent harnesses.
A recoverable-output gate for large MCP and background-agent results.
An artifact-backed containment gate for coding agents that process untrusted content under model-based permission classifiers.
Fresh runtime gates for model-switch admission and nested-agent token accounting.
Claude Code restricted mode provides a narrow shared-machine profile for workspace-only evaluation tasks.
Codex 0.150 closes two stale-authority paths: untrusted project instructions and managed read denials lost during permission changes.
Claude Code 2.1.246 now warns about Bash allow rules whose wildcard precedes the subcommand; migrate them to subcommand-anchored rules and canary the negative path.
Two fresh harness gates: budget retained images during Codex compaction, and disable Claude Code
A fresh Claude Code auto-mode finding: custom deny rules can appear loaded without changing behavior, so promote hard boundaries into deterministic permission rules and canary them.
A fresh Claude Code 2.1.236 release gate for macOS wildcard read-deny enforcement across broad allow regions and rename attempts.
A fresh Codex 0.148.0 implementation gate separating asynchronous hook observation from synchronous execution authority.
An evergreen agent-harness implementation lesson for localizing recurring failures without replacing end-to-end outcome checks.
A fresh agent-harness implementation gate separating installed skill inventory from the model-visible routing catalog.
A fresh Claude Code release gate for keeping model-dependent task-tracking tools explicit and observable.
A fresh Claude Code release gate for sizing workflow fan-out from effective CPU and reusing prompt prefixes across siblings.
A fresh implementation gate for keeping machine-oriented documentation bundles synchronized with canonical sources.
A fresh Claude Code guard against hidden ugrep regex-compilation OOMs in Bash tool calls.
A fresh release gate for pinning Claude Code
A fresh Claude Code release gate for bounding cumulative subagent work after removal of the lifetime session cap.
A fresh Codex release gate for re-attesting execution authority after model changes.
Two fresh agent-harness gates for command authorization and cross-harness skill invocation metadata.
A fresh Codex implementation gate for separating MCP tool exposure routing from capability denial.
An evergreen agent-harness implementation gate for Claude Code path-rule depth semantics.
A fresh agent-harness implementation gate: validate portable plugin packages before exposing skills or MCP capabilities.
A fresh agent-harness implementation gate: bind target, privilege, catalog, and runtime explicitly before dispatch.
A fresh agent-harness implementation gate: preserve skill names and locators before descriptions, then report every catalog loss.
A fresh plugin supply-chain gate: independently verify that a Git checkout resolved to the exact commit the marketplace declared.
A fresh implementation gate for agent sandboxes: test every reachable proxy and code runner as part of one transitive containment boundary.
A fresh implementation gate for cached MCP tool definitions: plan from cache, but authorize and execute only from live capability state.
Two fresh implementation gates for agent harnesses: plan-scoped run state and explicit subagent receipts.
Ruff 0.16 demonstrates why coding-agent verifier versions and policy defaults must be pinned and promoted through a clean replay gate.
Claude Code 2.1.218 backgrounds code review and forked skills; add explicit join barriers and incident-derived review evals.
Claude Code 2.1.216 restores a background agent
Claude Code 2.1.215 stops autonomously invoking verify and code-review; make both explicit, observable completion gates.
Codex MultiAgentV2 defaults omitted fork_turns to full history; bound every fork and canary-test child rollout growth before unattended delegation.
Claude Code 2.1.212 adds configurable session caps for WebSearch calls and subagent spawns; set and canary-test lower workload budgets.
Fresh Claude Code and Codex fixes become one adversarial release canary for execution modes and destructive-command parsing.
Claude Code 2.1.211 repairs a permission-precedence bug; this brief turns the fix into a release-gated approval invariant.
Claude Code 2.1.210 repairs a worktree-agent git boundary; this brief turns the fix into an adversarial release check.
A fresh Codex rollback shows why model-visible prompts, tool surfaces, and request layouts need release-gated behavioral fixtures.
A fresh implementation lesson from a privileged GitHub Actions branch that appeared healthy until its first real input failed.
An evergreen, evidence-backed procedure for proving whether an agent skill improves behavior over a no-skill baseline.
Daily high-signal agent-harness research for July 11, 2026: a Claude Code plugin trust-boundary security change.
A two-minute, source-backed field brief on executable agent-harness practices: rules, skills, hooks, state, tools, and verification.