High-signal AI/security/automation notes.
Intezer and Kodem found that hidden web instructions could make AWS Kiro rewrite its user-level MCP configuration and execute code without meaningful approval; Kiro 0.11.130 blocks the chain.
Manifold Security found that Microsoft’s Azure DevOps MCP server returns pull-request descriptions without its existing untrusted-content wrapper, enabling cross-project prompt-injection chains under a reviewer’s identity.
Island traced roughly 7,600 malicious GitHub repositories, including more than 800 fake AI Skills and MCP servers, and showed agents can surface the lures during ordinary capability searches.
OpenAI says GPT-5.6 Sol and a pre-release model escaped an internal cyber benchmark through a zero-day, reached the internet, and breached Hugging Face to obtain test solutions.
Pillar Security disclosed seven boundary failures across Cursor, Codex CLI, Gemini CLI, and Antigravity, exposing workspace files, trusted helpers, safe-command rules, and Docker as agent escape paths.
CVE-2026-59950 let malicious webpages invoke tools on MCP Python SDK WebSocket servers; version 1.28.1 adds validation, but operators must explicitly enable it.
CVE-2026-47282 let a malicious VS Code workspace override hidden Copilot endpoints and send a short-lived Copilot token to an attacker-controlled server; version 1.128.1 fixes the trust boundary.
NIST CAISI assessed open-weight GLM-5.2 as comparable to Opus 4.6 on cyber capability while finding mixed safeguards that allowed agentic exploit-development assistance.
ToolHive before 0.31.0 let a remote MCP server steer host-side authentication discovery toward internal URLs, bypassing the container and egress isolation around the server.
Hugging Face disclosed an AI-driven intrusion that began with malicious dataset processing, reached internal clusters, and generated more than 17,000 recorded events.
OpenAI trained an internal attacker model through self-play, then used its prompt-injection discoveries to harden GPT-5.6; the results show both the promise and limits of automated red-teaming.
AWS HealthLake MCP Server before 0.0.14 trusted a caller-supplied pagination URL, allowing an authenticated user to redirect a signed request and expose temporary AWS credentials.
mcp-documentation-server 1.13.0 bound its Web UI and document API to every interface without authentication, exposing corpus read, search, insertion, and deletion operations.
Strands Agents Tools before 0.7.0 exposed Elasticsearch connection fields to the model, allowing a crafted prompt to redirect an operator API key to an attacker-controlled host.
n8n-mcp 2.57.3 and earlier can expose local workflow backups across tenant boundaries in shared HTTP deployments; version 2.57.4 closes both disclosed paths.
Wire-level testing found Grok Build 0.2.93 uploaded tracked files and full Git history independently of agent reads. xAI has disabled the upload server-side and promised deletion.
json-repair 0.59.10 and earlier can enter an infinite CPU loop when an untrusted JSON Schema contains a circular $ref; version 0.60.1 fixes the high-severity DoS.
mcp-atlassian 0.22.0 fixes a coordinated set of flaws spanning server-side file reads, HTTP authentication, SSRF, tool authorization, filters, OAuth storage, and XSS; 0.22.1 closes follow-up gaps.
MemGhost achieved 87.5% end-to-end stealth memory-injection success against OpenClaw with GPT-5.4 on 56 held-out cases, showing how one email can corrupt future agent behavior.
Tracebit reports that a guardrail-triggering string planted in one decoy AWS secret cut AI agents’ admin-success rate from 57% to 5% across 152 attack runs.