High-signal AI/security/automation notes.
New research demonstrates that attacker-controlled log fields (URLs, user agents, payloads) can carry prompt injection attacks into LLM-based security operations center tools, with persona hijacks bypassing 68% of triage alerts.
IRGC-affiliated Nimbus Manticore group deploys AI-assisted malware toolkit MiniFast targeting defense, aerospace, and telecom sectors.
Three chained vulnerabilities in Claude.ai — invisible prompt injection via URL parameters, Files API data exfiltration, and an open redirect — create a complete attack pipeline against millions of users.
Phishing campaigns use zero-font and color-matched hidden text to inject benign content that tricks AI email filters into classifying malicious messages as safe.
PromptArmor demonstrated that weaponized workflow content can force Copilot Cowork to generate and share Microsoft 365 file links to unauthorized recipients, completely bypassing user consent.
A single-character Host header injection in Starlette (325M weekly downloads) bypasses path-based auth in FastAPI, vLLM, LiteLLM, and MCP servers, exposing credentials and PII across millions of AI agent deployments.
A new arXiv paper proposes mcp-attested: a signed clearance assertion, per-server tool allowlist, and tamper-evident audit log for MCP deployments.
INFRASCOPE uses reference-driven multi-agent analysis to find recurring vulnerability patterns across 688 AI infrastructure repositories, uncovering 20+ new vulnerabilities.
First measurement study of 7,973 live remote MCP servers finds over 40% expose tools without authentication; OAuth deployments universally exhibit at least one flaw, yielding 9 CVEs.
A new paper shows adversarial log content can inject into LLM-based SOC analyst assistants via persona hijacks and context manipulation — with 96% success on summarization tasks.
SUDP formalizes how LLM-driven agents can exercise user-authorized operations without gaining reusable authority over secrets, solving the Agent Secret Use problem.
Researchers from Google, UCSD, and UW-Madison argue enterprises must treat AI models as fundamentally untrusted and enforce security at the system level around agents, not inside the model.
Viper-MCP, the first end-to-end automated vulnerability auditing framework for MCP servers, scanned 39,884 repositories and discovered 106 confirmed 0-days, with 67 CVEs assigned to date.
New AudioHijack research shows imperceptible audio perturbations can inject malicious commands into voice AI assistants, achieving up to 96% attack success across 13 models.
The Consensus MCP server embeds hidden promotional text inside Claude tool instructions, forcing the model to pitch premium subscriptions — exposing a structural MCP supply-chain vulnerability.
Mitiga Labs demonstrates how a legitimate-looking AI agent skill can silently exfiltrate an entire codebase via poisoned Definition of Done instructions.
A fresh arXiv preprint proposes AgentWall, an OS-level safety layer that intercepts AI agent shell commands, API calls, and file writes before execution.
An arXiv paper evaluates GNN-based and classical detectors for attack detection over MCP tool-call traffic, finding content embeddings essential and naive random-split evaluations misleading.