self.md radar — 2026-07-08
agents got shoved into three unglamorous boundaries today: who may fetch the web, where memory sits, and what counts as trusted data.
start with robots.txt, because the old file is suddenly a receipt for the wrong war. then memory papers drag retrieval into the agent loop. finally, an agent-security paper names the ugly bit: a lot of agents still cannot tell trusted control data from hostile stuff they just scraped off the floor.
1. robots.txt kept guarding the training panic
sources:
what happened: Sitedex read the robots.txt files of the top 10,000 sites and found a very 2023-shaped scar. 38% of the GPTBot rules it could date were written in a single quarter after GPTBot arrived and the New York Times sued OpenAI. even stranger: 87% of those panic-window sites later came back and edited their files, but the new lines mostly kept naming training crawlers. ClaudeBot shows up on 671 of these sites, PerplexityBot on 540, while the answer-time fetchers — OAI-SearchBot, ChatGPT-User, Claude-User, Claude-SearchBot, Perplexity-User — get much less explicit treatment.
why this matters: Publishers are still arguing with the crawler that trains the model, while the crawler that answers the live question often walks through a different door. for self.md, that turns machine-readable publishing into something sharper than SEO hygiene: decide who may train, who may fetch, and who gets to quote you when the user never clicks.
2. agent memory hit the placement problem
sources:
what happened: Two new arXiv papers put numbers under the thing everyone hand-waves as “memory.” Memory in the Loop studies retrieval that is read and written on every reasoning step, not once per turn; the paper says expensive networked retrieval can inflate end-to-end latency by up to 83x, while an in-process store can answer around 100 microseconds. Akashic attacks the serving side with MemAttention, bounded chunks, and storage placement; across its workloads it reports up to 10.2 accuracy points, 1.21x throughput, and 1.88x sustainable request rate over prior memory baselines.
why this matters: Memory is becoming part of the hot path, not a polite second-brain drawer the agent checks before work. once retrieval sits inside the loop, architecture matters: where the memory lives, how often it is touched, and whether the system can afford remembering without turning every step into molasses.
3. agent data injection found the trust seam
sources:
what happened: Agent Data Injection attacks are a nastier cousin of indirect prompt injection: malicious data is disguised as trusted metadata, resource identifiers, origins, tool output, or context. the paper reports arbitrary-click attacks against Claude in Chrome, Antigravity, and Nanobrowser, plus remote-code-execution and supply-chain attacks against Claude Code, Codex, and Gemini CLI. CLRK landed in the same pressure zone from the builder side: a Kubernetes-native agent runtime that runs workloads in gVisor sandboxes and intercepts LLM APIs, MCP traffic, and tool calls at the egress boundary.
why this matters: “don’t obey hostile instructions” is too small a security model for agents that browse, click, clone, install, and run. the real seam is data/control separation: the agent needs to know which bytes are merely evidence and which bytes are allowed to steer the machine.
supporting links
- Rowboat — a local-first desktop AI coworker with markdown storage, a knowledge graph, browser, notes, code mode, and background agents; useful proof that the chat app is turning into a work surface.
- OfficeCLI — Word, Excel, and PowerPoint get a CLI/MCP surface for agents, including render-and-fix loops and resident mode for multi-step document work.
- TencentDB Agent Memory — local long-term agent memory with a four-tier pipeline; support-grade, but it rhymes with the memory-placement papers.
- Anthropic’s global workspace work — interpretability research on what reaches Claude’s reportable workspace; not an ops launch, but adjacent to today’s memory/control question.
left on the table
- agent-skills stayed out because it was already used on 2026-07-07. exact repeats are how a radar turns into a soup ladle.
- page-agent was also a recent repeat from 2026-07-06, and today’s browser/action story had stronger security evidence.
- CubeSandbox had sandbox fit, but the sandbox/runtime lane was just covered on 2026-07-04; CLRK only appears here because it supports the ADI trust-boundary story.
- memory is a repeat theme, allowed today because the object changed: these papers are about loop-time latency and serving placement, not generic “agents need memory.”
Related self.md routes
- Personal AI OS tools — the control-plane map for personal agents, receipts, memory, and tools
- AI coding assistants — compare coding workbenches by review surface, permissions, cost, logs, and escape hatches
- Best Claude Code plugins — choose the Claude-specific extensions worth installing, and the ones to skip