self.md radar — 2026-07-13
agent runs started showing their meter before the work even began: hidden harness tokens, hour-long terminal attempts, and repo memory that only pays off when it can prove what it saved.
the useful through-line is cost with receipts. Systima measured what two coding harnesses send before the user prompt. Long-Horizon-Terminal-Bench made terminal agents sit with multi-hour work instead of toy tasks. capn-hook is the practical counter-move: stop paying a fresh agent to rediscover the same code paths every morning.
1. the prompt arrived after the bill
sources:
what happened: Systima put Claude Code and OpenCode on the same Anthropic models, same machine, and same tasks, then logged the requests between the coding harnesses and Anthropic’s endpoint. in the smallest test, a one-line reply on Sonnet 4.5, Claude Code sent roughly 33,000 tokens of system prompt, tool schemas, and scaffolding before the user prompt; OpenCode sent about 7,000. the gap narrowed on Claude Fable 5 to about 3.3x, which is almost the more useful finding: the “agent” cost is not just model price, it is harness design, cache stability, instruction files, MCP servers, subagents, and whatever gets injected before you type.
why this matters: agent pricing arguments keep pretending the model is the invoice. the harness is quietly becoming the meter reader, and sometimes it starts billing before the doorbell rings.
2. long terminal work got a harsher benchmark
sources:
what happened: Long-Horizon-Terminal-Bench introduced 46 terminal tasks across nine categories, including experiment reproduction, software engineering, multimodal analysis, games, and scientific computing. unlike short terminal benchmarks that only grade the final outcome, these tasks are decomposed into fine-grained subtasks so partial progress can be scored. the authors evaluated 15 frontier models and report an average of 9.9 million tokens, 231 episodes, and 85.3 minutes per run. even the strongest tested model reached 15.2% pass@1 at a 0.95 partial-reward threshold and 10.9% at perfect reward.
why this matters: this is the shape real agent work keeps taking: not one clean solve, but a long, expensive scrape across state, logs, retries, and partial credit. if your eval ends before the agent gets tired, it is probably grading the demo, not the job.
3. codebase memory got a token receipt
sources:
what happened:
capn-hook shipped a small memory loop for coding agents: before searching a repo, the agent can ask capn; after an expensive discovery, it charts the files that answer the question. each saved answer is backed by file hashes, and stale charts delete themselves when those files change. the project’s eval ran 60 developer questions across Dub, Polar, PostHog, Twenty, and Documenso, with baseline, chart, and recall arms for 180 runs. recall used 77.5% fewer tokens than cold exploration, with 60/60 correct answers in both baseline and recall, and a break-even of about 1.6 recalls.
why this matters: memory usually arrives dressed as a personality feature. this one is more useful because it is boring: a local map of “where does this live?” that expires when the coastline moves.
supporting links
- Sanbox — isolated, resumable microVM runs for agents, with tool calls, model activity, artifacts, agent state, and conversation history preserved for later inspection.
- Open-Inspect background agents — open-source background coding agents with Slack, GitHub, Linear, cron, webhooks, PR creation, and a very honest single-tenant security warning around shared GitHub App access.
- Baton — a macOS menu-bar watcher for Claude Code and Codex sessions, built around the small operational problem of knowing which agent is waiting on you now.
- wigolo — local-first MCP web search, fetch, crawl, extract, cache, and research loops with no API key or metered cloud bill.
left on the table
- planning-with-files had big GitHub heat and fresh v3 benchmarking, but its URL was already in the recent seen ledger. useful, not today’s main.
- Adaptive Recall , exxperts , and Loci all fit the memory wave, but today’s main memory slot went to the one with a repo-question eval and hash-based invalidation.
- Sanbox and Microsoft MXC were tempting runtime-boundary candidates, but sandboxing has been covered hard this week. they work better as supporting pressure than another main signal.
- home-assistant/core , anthropics/claude-cookbooks , and the usual trending repos were traffic, not signal.
Related self.md routes
- Personal AI OS tools — the control-plane map for personal agents, receipts, memory, and tools
- AI coding assistants — compare coding workbenches by review surface, permissions, cost, logs, and escape hatches
- Best MCP servers — connect files, browsers, memory, search, and workflow tools without turning the stack into soup