High-signal AI/security/automation notes.
A new paper shows adversarial log content can inject into LLM-based SOC analyst assistants via persona hijacks and context manipulation — with 96% success on summarization tasks.
SUDP formalizes how LLM-driven agents can exercise user-authorized operations without gaining reusable authority over secrets, solving the Agent Secret Use problem.
Researchers from Google, UCSD, and UW-Madison argue enterprises must treat AI models as fundamentally untrusted and enforce security at the system level around agents, not inside the model.
Viper-MCP, the first end-to-end automated vulnerability auditing framework for MCP servers, scanned 39,884 repositories and discovered 106 confirmed 0-days, with 67 CVEs assigned to date.
New AudioHijack research shows imperceptible audio perturbations can inject malicious commands into voice AI assistants, achieving up to 96% attack success across 13 models.
The Consensus MCP server embeds hidden promotional text inside Claude tool instructions, forcing the model to pitch premium subscriptions — exposing a structural MCP supply-chain vulnerability.
Mitiga Labs demonstrates how a legitimate-looking AI agent skill can silently exfiltrate an entire codebase via poisoned Definition of Done instructions.
A fresh arXiv preprint proposes AgentWall, an OS-level safety layer that intercepts AI agent shell commands, API calls, and file writes before execution.
An arXiv paper evaluates GNN-based and classical detectors for attack detection over MCP tool-call traffic, finding content embeddings essential and naive random-split evaluations misleading.
Microsoft open-sources an AI Agent Governance Toolkit covering all 10 OWASP Agentic Top 10 risks, with policy enforcement, zero-trust identity, and execution sandboxing for autonomous agents.
Two critical Microsoft Semantic Kernel flaws (CVSS 9.8) escalate indirect prompt injection into remote code execution on agent hosts via eval() and unsafe file download.
An ACM paper testing 10 open-source models across 167 attack scenarios finds multi-turn reasoning jailbreaks remain unsolved even with lightweight defenses applied.
Three prompt injection variants that bypass traditional WAF and EDR detection: indirect RAG injection, second-order tool-call injection, and conversation-history poisoning.
TrapDoor campaign weaponizes .cursorrules and CLAUDE.md files via cross-ecosystem supply chain attacks on npm, PyPI, and Crates.io to hijack AI coding assistants.
1Password partners with OpenAI to deliver an Environments MCP Server for Codex, issuing just-in-time scoped credentials that never appear in prompts or model context.
Adaptive Security research reveals an 8-to-1 gap between shadow AI adoption and enterprise governance, with OAuth scopes representing the real data-exposure surface.
New cryptographic protocol binds credential validity to parent liveness proofs, achieving 90× faster revocation than OAuth 2.0 for AI agent swarms.