Skip to content

[!] TOPIC ARCHIVE // #AGENT-SECURITY

#Agent-Security

All 19 guides, operator dossiers, and signals tagged with #Agent-Security.

RADAR SIGNAL

the agent’s paper trail

A security patch, fresh-context agents, confidence probes, and the interface of a useful agent reply.

Read →
RADAR SIGNAL

the accidental message board

OpenAI’s Black Hat reconstruction shows agents using writable Artifactory files and a WebDAV cache as accidental cross-run communication surfaces during the Hugging Face incident.

Read →
RADAR SIGNAL

eval sandboxes, compiler policy, review queues

Anthropic found Claude eval runs touching real systems, GCC drew a line around LLM-generated contributions, and GitHub shipped stacked PRs for the review bottleneck.

Read →
RADAR SIGNAL

documents, sandboxes, cryptography bills

Word documents can carry hidden agent instructions, eval sandboxes can become intrusion launchpads, and cryptography work now arrives with API bills, artifacts, and disclosure receipts.

Read →
RADAR SIGNAL

breaches, DAGs, counters

a cyber benchmark reached Hugging Face production, enterprise analysis got compiled into DAGs, and work agents moved behind queues, sandboxes, approvals, and receipts.

Read →
RADAR SIGNAL

full access, endpoint maps, tuning rights

Codex exposed the filesystem blast radius of full-access agents, Traceforce mapped AI apps on devices, and Kimi plus Inkling moved open weights toward customization.

Read →
RADAR SIGNAL

fetch leaks, loop brakes, agent APIs

A Claude memory leak exposed browsing as egress policy, LoopGain measured when agent loops should stop, and new API work treated agents as first-class callers.

Read →
RADAR SIGNAL

consent gates, hard walls, wrapper leaks

Samsung tied health sync to AI-training consent, coding-agent tools moved trust into VMs and effect systems, and new eval papers showed wrappers and relays can change the result.

Read →
RADAR SIGNAL

receipts before trust

Grok Build gets a wire-level receipt, MCP servers get a security scorecard, and personal agents get signed mandates and succession files.

Read →
RADAR SIGNAL

the agent needs a human on the hook

Godot tightened AI contribution rules, Senior SWE-Bench graded agents on real senior-engineer work, and dependency checks moved into agent preflight.

Read →
RADAR SIGNAL

hidden marks, owned memory, parser risk

Claude Code prompt marks, host-owned agent memory, and a protobuf parser bug pushed today’s control story into the small layers.

Read →
RADAR SIGNAL

setup chains got teeth

agent work is getting squeezed from both sides: attackers are learning the setup path, benchmarks are exposing harness blur, and runtimes are turning machine management into product surface.

Read →
RADAR SIGNAL

outside hands

model access, bug telemetry, and agent work logs all showed the same pressure: personal AI systems need receipts when outside hands can steer the machine.

Read →
RADAR SIGNAL

permission has to leave the prompt

Meta’s AI support path, PyTorch’s coding-agent policy, and JetBrains Mellum2 all point at the same move: agents need real permissions, owners, and handoff logs.

Read →
RADAR SIGNAL

trust surfaces

citation tables, permission prompts, and desktop drivers all turned into trust surfaces today: EY’s fake references, Anthropic’s containment math, and Cua’s installable computer-use layer.

Read →
RADAR SIGNAL

agent job control

Agents crossed into job control: Codex Goals and agent spend got budgets, CUDA verifier wins failed workload tests, and AGENTS.md changes needed holdouts.

Read →
RADAR SIGNAL

permission surfaces

Copilot Cowork exposed the approval gap, llama.cpp patched local-agent checkpoints, and Quartz v5 turned Obsidian publishing into plugin infrastructure.

Read →
RADAR SIGNAL

May 25 Radar: serial reviewers, memory bills, and leaking sandboxes

Addy Osmani and Armin Ronacher put a human bottleneck under coding agents, Epoch and an energy paper put numbers under agent costs, and Canister shows why network allow-lists still leak secrets.

Read →
ENTRY

ClawSec

complete security skill suite for OpenClaw agents — drift detection, skill integrity, automated audits, SOUL.md protection

Read →
← All topics & tags