self.md radar — 2026-10-02
the agent’s name is not the whole bill
A benchmark, a desktop-memory app and a research harness land on different parts of the same working mess: what happened, who was allowed to do it, and how much extra machinery was involved.
1. A coding agent’s harness can double the bill
Arena tested 21 model–harness pairs across SWE-bench Lite and Terminal-Bench 2.0. In one comparison, Claude Fable 5 solved 97.8% of SWE-bench Lite attempts in Claude Code and 96.7% in Pi. The reported cost was about $1.33 in Claude Code and $0.67 in Pi. The paper used 30 sampled tasks from each benchmark, three runs per pair and a 100-turn cap.
That does not settle which harness to use. It does spoil the lazy comparison where the model name carries the whole result. Tool defaults, instructions and execution loops can add a large cost without buying much more on the task set in front of you.
reading: Arena, “HarnessTax: How Much Does the Harness Matter for Coding Agents?”
2. Breadcrumb puts exclusions where the memory pitch usually gets foggy
Breadcrumb is a Mac app for feeding prior screen activity, meetings, day summaries and AI sessions back into an assistant. Its page says you can exclude apps, sites, words and phrases; incognito sites are never recorded; and pausing stops recording. History is encrypted at rest, with the key held in Keychain. The app says it needs no account, email, subscription or telemetry.
Those are product claims, not an audit. Still, they give the reader something better than “private by design”: a set of switches and custody statements to inspect. Its own small print also keeps the hosted-model boundary visible. Ask Claude or ChatGPT and the answer goes to that service; use a local model and it stays local.
reading: Innerloop, “Breadcrumb”
3. Praxa separates a proposal from an effect
The Praxa paper treats proposal, authority, dispatch, verified external effect and promotion as different states. Its Terminal-Bench Core pilot gave both the baseline and its reliability layer 17 passes from 36 strict trials. The reliability layer used 37.49% more input tokens and 50.73% more output tokens, so the authors do not claim a performance win.
That restraint is more useful than another agent victory lap. A tool call can succeed while the external record does not change. For work that touches a repository, calendar, payment or database, the read-back is not a decorative receipt. It is the first point at which “done” can mean anything.
Related self.md routes
- Personal AI OS tools — the control-plane map for personal agents, receipts, memory, and tools
- AI coding assistants — compare coding workbenches by review surface, permissions, cost, logs, and escape hatches
- Local LLM runtimes — pick local inference by custody, speed, privacy, cost, and failure mode