self.md radar — 2026-07-12
agent trust had a paperwork day: the good stuff was not a smarter demo, it was a receipt, a scorecard, and a signed boundary.
start with the Grok Build teardown, because that is the one with teeth. then the MCP server index turns tool trust into something procurement can actually reject. the wildcard is tiny but interesting: a personal-agent spec that treats mandates and succession as signed objects, not vibes with an API key.
1. Grok Build got traced at the wire
sources:
what happened:
A GitHub gist published a wire-level teardown of xAI’s official Grok Build CLI, testing grok 0.2.93 on macOS with captured traffic and throwaway canary secrets. the author says Grok sent file contents, including a .env, through POST /v1/responses, then also uploaded a session_state archive through POST /v1/storage. the sharper bit: with the prompt “reply OK, do not read any files,” the tool allegedly uploaded the whole repo as a git bundle anyway; in a 12 GB repo test, /v1/storage moved 5.10 GiB while the model-turn channel moved 192 KB. the storage path is tied in the writeup to a Google Cloud Storage bucket named grok-code-session-traces, and the author is careful about the line they did not prove: transmission and storage, yes; training use, no.
why this matters: Native coding CLIs are becoming data-movement systems with a chat box attached. if the repo can leave as a bundle, “the agent may read this directory” is a child lock on a warehouse door.
2. MCP servers got a scorecard before the agent connects
sources:
what happened: Canopii launched a public Trust Index for MCP servers, pulling from the official MCP registry and scanning server versions for tool poisoning, prompt injection, supply-chain, dependency, credential, transport, auth, and maintenance risk. its homepage reports 16,129 servers, 12,629 scored, 4,263 live-verified endpoints, and 1,262 high-risk D/F entries. the methodology is blunt in the right way: confirmed flaws cap the score, low scan confidence caps the score, and popularity only gets a small nudge. a 10k-star server can still be a bad idea.
why this matters: MCP made tool installation feel like adding Lego bricks to an agent. this is the counterweight: an allow-list can be backed by evidence instead of “seems fine, shipped by someone with a nice README.”
3. personal agents got a succession file
sources:
what happened: macrokit/agent-world published a draft standard, SDK, and live hub for personally-owned agents. it is tiny, yes, but unusually explicit: agents have Ed25519 identifiers, canonical JSON signatures, signed manifests, capabilities, mandates, succession blocks, and portability rules. the spec’s “boundary-only” rule says the standard must not dictate the model, runtime, framework, or storage inside the agent. its longevity test is even stranger and better: every feature is judged against an orphaned-agent case, where the person is gone and the agent still needs to carry its frame across 30+ years of platform churn.
why this matters: Most agent products still treat ownership as account settings plus billing. this one asks the nastier question: what exactly is the agent allowed to commit for you, and who inherits that permission when you are not there to click approve?
supporting links
- Confessor — local forensic reports for Claude Code-style sessions: files opened, commands run, secrets seen, and sensitive-read-then-network-call chains.
- mindwalk — turns Claude Code and Codex session logs into a local 3D map of what the agent searched, read, edited, and verified.
- Aether — runs Claude Code, Codex, and OpenCode inside watchable cloud devboxes, with PR review, Slack/Linear handoff, and per-second metering.
- Google Stitch Skills — a big-vendor skills library for Stitch MCP, Codex, Gemini CLI, Claude Code, Cursor, and Antigravity; useful, but too close to yesterday’s skills thread for a main slot.
left on the table
- Google Stitch Skills had the GitHub heat, but yesterday already covered skills/governance as the prompt-to-procedure move. same story, bigger logo.
- Confessor , mindwalk , and Aether were good operator tools, but session replay and receipts were already near yesterday’s memory signal, so they belong as supporting evidence.
- Hamel’s automated evals post stayed out because the RSS item currently resolves to the blog root, not a stable article page, and eval/harness coverage has been loud all week.
- home-assistant/core , hashicorp/terraform , and the usual trending repos were traffic, not today’s signal.
Related self.md routes
- Personal AI OS tools — the control-plane map for personal agents, receipts, memory, and tools
- AI coding assistants — compare coding workbenches by review surface, permissions, cost, logs, and escape hatches
- Best Claude Code plugins — choose the Claude-specific extensions worth installing, and the ones to skip