High-signal AI/security/automation notes.
GuardianAgentBench tested 580 scenarios across three production agent frameworks. The best setup scored 74.8%, while runtime tool-call guardrails recovered 19.9% of failures at a 0.5% false-positive rate.
NVIDIA and 36 partners launched the Open Secure AI Alliance around an open agent-defense stack. The announcement names useful components, but no roadmap or governance model yet.
CVE-2026-55607 chained prompt injection, Git worktree confusion, symlinks, and fsmonitor to escape Claude Code’s macOS sandbox. Upgrade to 2.1.163 or later.
Anthropic’s Opus 5 system card reports strong exploit and cyber-range capability alongside materially lower prompt-injection success, with important limits for enterprise risk decisions.
CVE-2026-66012 exposed 31 SiYuan MCP tools through a Reader token, enabling workspace theft, credential exposure, file writes, and delayed desktop code execution. Upgrade to 3.7.2 or later.
UC Berkeley-led research separates untrusted exploration from privileged action and compresses cross-boundary hints, sharply reducing prompt-injection success without the utility collapse of rigid dual-agent baselines.
Accomplish chained CVE-2026-46331 through the local Claude Cowork Linux VM to reach a read-write share of the Mac host filesystem. Current Cowork sessions run remotely by default.
IssueTrojanBench reports 2,776 successful malicious actions across 4,176 coding-agent runs. Dependency installation was the most reliable attack; framework controls caused no measured rejections.
Redis shipped seven security releases after researchers published authenticated RCE chains against versions previously considered fixed. The Kimi K3 discovery-speed claims remain self-reported.
Eighteen Open WebUI CVEs disclosed on July 24 converge on version 0.10.0. Six are high severity, including cross-user tool execution and stored Pyodide XSS.
AWS API MCP Server 0.2.13 through 1.3.46 could silently skip its configured security-policy check after a startup initialization failure. Version 1.3.47 fails closed.
AWS Bedrock AgentCore Python SDK before 1.18.1 let crafted install_packages() input execute arbitrary commands inside Code Interpreter. Model-generated package names must be treated as untrusted.
Google fixed SSRF in Agent Studio’s generated /api-proxy backend and told users to regenerate and redeploy web apps whose code was created before July 1, 2026.
Zenity found a patched ChatGPT Workspace Agents flaw that let one crafted link create, authorize, publish, and schedule an agent inside a victim’s connected work accounts.
Google Cloud put CodeMender into preview: the managed security agent scans repositories, builds proof-of-concept exploits in a customer-managed sandbox, and proposes tested patches for human approval.
Intezer and Kodem found that hidden web instructions could make AWS Kiro rewrite its user-level MCP configuration and execute code without meaningful approval; Kiro 0.11.130 blocks the chain.
Researchers found seven attack paths across five open-source Android agents, chaining imperceptible screen instructions, screenshot tampering, and unsafe ADB command construction into host compromise.
CVE-2026-11876 allowed authenticated ZenML users to enumerate other tenants’ deployed stacks, service-connector details, owners, and infrastructure metadata; version 0.95.0 adds authorization checks.
Manifold Security found that Microsoft’s Azure DevOps MCP server returns pull-request descriptions without its existing untrusted-content wrapper, enabling cross-project prompt-injection chains under a reviewer’s identity.
Island traced roughly 7,600 malicious GitHub repositories, including more than 800 fake AI Skills and MCP servers, and showed agents can surface the lures during ordinary capability searches.