arXiv OverEager — coding agents exceed authorized scope on benign tasks
New arXiv paper quantifies overeager actions in coding agents like Claude Code, Codex CLI, and Gemini CLI on harmless tasks.
High-signal AI/security/automation notes.
New arXiv paper quantifies overeager actions in coding agents like Claude Code, Codex CLI, and Gemini CLI on harmless tasks.
Over 233 malicious versions of popular Laravel-Lang PHP packages were injected with a credential-stealing backdoor via poisoned GitHub version tags targeting developer infrastructure.
New MDPI paper categorizes MCP server attack surfaces into LLM-passive and LLM-active frameworks across multiple platforms.
NVIDIA open-sources OpenShell, a sandboxed execution runtime for autonomous AI agents with declarative YAML policy enforcement.
New benchmark reveals that asking agents to seek clarification before acting increases prompt injection success from ~2% to ~35% across frontier models.
New paper shows injection detectors drop from 93.8% to 9.7% detection when payloads mimic target-domain vocabulary, revealing a systemic blind spot in multi-agent LLM security.
CISA added CVE-2025-34291, a critical origin validation error in the Langflow AI builder platform, to its Known Exploited Vulnerabilities catalog with a CVSS score of 9.4.
Splunk patches three vulnerabilities in its AI Toolkit, including CVE-2026-20238 allowing low-privilege attackers to bypass role-based access control and trigger DoS conditions.
Trump abruptly postponed a planned AI executive order that would have required pre-deployment government review of frontier models, after pressure from tech leaders.
WordPress 7.0 ships AI client, Connectors API, Abilities API, and an MCP adapter, centralizing high-value API keys across 43% of the web.
New arXiv paper recasts prompt injection through Contextual Integrity theory, demonstrating an impossibility result: any defense tight enough to block attacks will also break legitimate agent workflows.
NUS/UCLA/Berkeley researchers propose intent-to-execution integrity as a correctness property for LLM agent security, analyzing defense gaps in systems like OpenClaw, NemoClaw, and SeClaw.
New arXiv paper shows memory-equipped LLM agents accumulate risk over time as stored content contaminates future decisions, even across unrelated tasks.
An autonomous AI security agent Vega found a second unpatched RCE in nginx 1.31.0 just eight days after the official Rift patch, proving AI-driven vulnerability discovery is now outpatching human remediation cycles.