Microsoft Security Blog — Running OpenClaw safely
Microsoft’s security team outlines why self-hosted OpenClaw agents should be treated as untrusted code execution and run in isolated environments with dedicated credentials.
High-signal AI/security/automation notes.
Microsoft’s security team outlines why self-hosted OpenClaw agents should be treated as untrusted code execution and run in isolated environments with dedicated credentials.
NIST launches the AI Agent Standards Initiative, focusing on security, interoperability, and identity standards for agentic systems.
Check Point Research shows how web-based AI assistants with browsing/URL fetching can be abused as covert command-and-control relays.
mbgsec reports a prompt-injection attack on Cline’s Claude issue-triage workflow that stole npm publishing credentials and led to an unauthorized package release.
Google Threat Intelligence Group outlines rising model distillation attempts and deeper AI integration across attacker workflows.
Praetorian details how malicious MCP servers can chain trusted tools to exfiltrate data, execute code, and manipulate AI assistants.
Cerbos outlines why MCP servers need fine-grained authorization and highlights real incidents tied to weak access controls in agent tooling.
PromptArmor details how messaging-app link previews can leak data from AI agents by auto-fetching attacker-controlled URLs without user clicks, with OpenClaw + Telegram exposed by default.
Snyk argues AI agent security needs guardrails that intercept tool calls, inspect context, and enforce policy before actions execute.
Straiker STAR Labs reports a SmartLoader campaign that trojanized an Oura MCP server and poisoned MCP registries to deliver StealC infostealer payloads.
University of Toronto InfoSec outlines how MCP deployments amplify existing threats and how risk changes across local, org-hosted, multi-tenant, and third-party models.
A new paper proposes cryptographic provenance for prompts and context, aiming to make LLM workflows tamper-evident and policy-enforced.
Microsoft Security Blog outlines common Copilot Studio agent misconfigurations and the Defender hunting queries that can detect risky sharing, weak auth, and tool misuse.
OWASP GenAI Security Project publishes a practical guide for securing Model Context Protocol (MCP) servers with concrete architecture, auth, validation, isolation, and deployment controls.
A new paper analyzes hidden activations across GPT-J, LLaMA, Mistral, and Mamba to detect jailbreaks and disrupt them at inference time.
A new paper proposes autonomy metrics and a security-aware planning agent that improves HITL throughput under prompt-injection defenses.
Cyata reports three prompt-injection CVEs in Anthropic’s official mcp-server-git, enabling path bypass, argument injection, and destructive Git actions.
LayerX reports a zero-click RCE chain where a Google Calendar event can trigger Claude Desktop Extensions to execute local code via MCP connectors.