[!] TOPIC ARCHIVE // #AGENT-SECURITY
#Agent-Security
All 19 guides, operator dossiers, and signals tagged with #Agent-Security.
the agent’s paper trail
A security patch, fresh-context agents, confidence probes, and the interface of a useful agent reply.
the accidental message board
OpenAI’s Black Hat reconstruction shows agents using writable Artifactory files and a WebDAV cache as accidental cross-run communication surfaces during the Hugging Face incident.
eval sandboxes, compiler policy, review queues
Anthropic found Claude eval runs touching real systems, GCC drew a line around LLM-generated contributions, and GitHub shipped stacked PRs for the review bottleneck.
documents, sandboxes, cryptography bills
Word documents can carry hidden agent instructions, eval sandboxes can become intrusion launchpads, and cryptography work now arrives with API bills, artifacts, and disclosure receipts.
breaches, DAGs, counters
a cyber benchmark reached Hugging Face production, enterprise analysis got compiled into DAGs, and work agents moved behind queues, sandboxes, approvals, and receipts.
full access, endpoint maps, tuning rights
Codex exposed the filesystem blast radius of full-access agents, Traceforce mapped AI apps on devices, and Kimi plus Inkling moved open weights toward customization.
fetch leaks, loop brakes, agent APIs
A Claude memory leak exposed browsing as egress policy, LoopGain measured when agent loops should stop, and new API work treated agents as first-class callers.
consent gates, hard walls, wrapper leaks
Samsung tied health sync to AI-training consent, coding-agent tools moved trust into VMs and effect systems, and new eval papers showed wrappers and relays can change the result.
receipts before trust
Grok Build gets a wire-level receipt, MCP servers get a security scorecard, and personal agents get signed mandates and succession files.
the agent needs a human on the hook
Godot tightened AI contribution rules, Senior SWE-Bench graded agents on real senior-engineer work, and dependency checks moved into agent preflight.
hidden marks, owned memory, parser risk
Claude Code prompt marks, host-owned agent memory, and a protobuf parser bug pushed today’s control story into the small layers.
setup chains got teeth
agent work is getting squeezed from both sides: attackers are learning the setup path, benchmarks are exposing harness blur, and runtimes are turning machine management into product surface.
outside hands
model access, bug telemetry, and agent work logs all showed the same pressure: personal AI systems need receipts when outside hands can steer the machine.
permission has to leave the prompt
Meta’s AI support path, PyTorch’s coding-agent policy, and JetBrains Mellum2 all point at the same move: agents need real permissions, owners, and handoff logs.
trust surfaces
citation tables, permission prompts, and desktop drivers all turned into trust surfaces today: EY’s fake references, Anthropic’s containment math, and Cua’s installable computer-use layer.
agent job control
Agents crossed into job control: Codex Goals and agent spend got budgets, CUDA verifier wins failed workload tests, and AGENTS.md changes needed holdouts.
permission surfaces
Copilot Cowork exposed the approval gap, llama.cpp patched local-agent checkpoints, and Quartz v5 turned Obsidian publishing into plugin infrastructure.
May 25 Radar: serial reviewers, memory bills, and leaking sandboxes
Addy Osmani and Armin Ronacher put a human bottleneck under coding agents, Epoch and an energy paper put numbers under agent costs, and Canister shows why network allow-lists still leak secrets.
ClawSec
complete security skill suite for OpenClaw agents — drift detection, skill integrity, automated audits, SOUL.md protection