Posts
High-signal AI/security/automation notes.
IMA Attack — Collaborative-Adversarial Jailbreak Hits 89% Against Multi-Agent Systems
The IMA attack framework achieves 89% jailbreak success against MetaGPT, CrewAI, and other multi-agent systems by exploiting collaboration dynamics that amplify harm over single-agent baselines.
Tencent AI-Infra-Guard — Multi-Layer Agent Red Teaming Framework
Tencent Zhuque Lab releases AI-Infra-Guard, an open-source framework spanning infrastructure, protocol, agent, and model layers for comprehensive AI agent red teaming.
arXiv — SkillCloak Bypasses 90%+ of Agent Skill Scanners With Payload-Preserving Evasion
New research demonstrates that static skill scanners fail against adaptive evasion techniques, with SkillCloak bypassing over 90% of defenses while preserving malicious functionality.
AvePoint: 88% of Organizations Report AI Agent Security Breaches as Visibility Gaps Triple
T3MP3ST — Open-Source Framework Turns AI Coding Agents Into Autonomous Red Teamers
T3MP3ST orchestrates Claude Code, Codex, and Hermes agents through a recon-to-exploit kill chain with no additional API keys or cloud infrastructure.
CVE-2026-52830 — fast-mcp-telegram Path Traversal Bypasses Session Protection
Critical CVSS 9.4 path traversal in fast-mcp-telegram lets remote attackers bypass session controls and authenticate as the default legacy user via crafted Bearer tokens.
DataDome — AI Agent Identity Crisis Drives H1 2026 Threat Landscape
ICML 2026 — Prompt Injection as Role Confusion and CoT Forgery
ICML 2026 paper traces prompt injection to role confusion in LLMs and introduces CoT Forgery, a zero-shot attack achieving 60% success against frontier models.
LiteLLM Three-CVE RCE Chain — Default User to Full AI Gateway Compromise in Two Seconds
Unit 42 — Malicious OpenClaw Skills on ClawHub Delivered Infostealers and Crypto Fraud
Palo Alto Networks Unit 42 found five malicious OpenClaw skills on ClawHub that evaded VirusTotal and ClawScan, delivering macOS infostealers and running agentic financial fraud.
Alibaba Bans Claude Code Over Alleged Covert Environment Detection
Alibaba plans to ban Claude Code across workplaces after a Reddit user alleged the tool covertly detected Chinese AI labs and modified its system prompt.
arXiv — Capability Gates Are Not Authorization: Confused-Deputy Failures in LLM Agent Frameworks
New research audits LangChain, LlamaIndex, and Stripe Agent Toolkit for per-call authorization — all three ship capability gating but none ships deterministic fail-closed authorization by default.
MPBench Exposes Memory Poisoning Attacks in LLM Agents
New research identifies 9 structural vulnerabilities in LLM agent memory systems and introduces MPBench, showing aggressive memory-writing agents are more exploitable and existing prompt injection defenses fail to cover memory poisoning.
Red-Teaming the Agentic Red-Team: Offensive Security Agents Have Shared Design Flaws
First in-depth security analysis of offensive security agents reveals shared design flaws enabling API key exfiltration, persistent footholds, and operator machine compromise even inside sandboxes.
NVIDIA — Fine-Tuned 30B Model Outperforms Frontier LLMs at Agent Exploitation
Black Hat 2026 briefing reveals a purpose-trained 30B open-source model achieves 56% exploit success rate against AI agents at 70-125x lower cost than frontier models.
Flowise CVE-2026-40933 — MCP stdio Transport RCE on Import
Obsidian Security discloses critical Flowise flaw where importing a malicious workflow file triggers remote code execution via MCP stdio transport.
Trend Micro — CVE-2026-33017 Langflow RCE Weaponized for Monero Mining on AI Servers
Threat actors exploit a CVSS 9.3 Langflow vulnerability to deploy custom Monero miners, turning exposed AI app endpoints into cryptojacking gateways.
LiteLLM PyPI Supply Chain Attack — Credential Stealer in Versions 1.82.7–1.82.8
LiteLLM PyPI package compromised — malicious versions 1.82.7 and 1.82.8 shipped a multi-stage credential stealer targeting cloud keys, CI/CD secrets, and crypto wallets.