self.md radar — 2026-10-04
The controls are catching up with the agent
The week’s agent news has a duller centre than the launch cycle suggests: someone has to choose the ceiling, keep the record, and decide which browser session is actually in play.
1. Put the spending ceiling in before the coding agent starts wandering
Simon Willison argues for default hard budget caps on coding-agent runs, with an explicit opt-in when a task genuinely needs a larger allowance. That puts a stopping condition inside the run, not in the invoice that arrives afterwards.
Usage dashboards are receipts. A ceiling is an operating decision. The useful default is not “let the agent finish”; it is “let it work inside a number somebody chose.” Otherwise autonomy is just an unattended meter.
reading: Simon Willison, “Default hard budget caps for coding agents”
2. A final diff is not a record of an agent run
agent-blackbox is an open-source project for recording and auditing AI-agent activity. Its premise is simple: the run itself should remain inspectable after an agent has touched files and called tools.
A clean diff cannot show the detours, tool calls, or first bad decision that produced it. That gap gets expensive when the agent moves faster than a person can replay terminal scrollback. A readable trace is not paperwork glued onto autonomy. It is part of the control surface.
reading: developerfred/agent-blackbox
3. “Give the agent the internet” is a pile of different permissions
Agent Reach packages web access for agents across services including GitHub, X, Reddit and YouTube. Its README separates public reading from flows that need login, cookies, or an existing browser session.
Browser access gets sold as one capability. In practice it is a corridor full of different doors: public pages, authenticated accounts, imported cookies, local sessions, and commands that can change a machine. Each service and session needs its own decision, not a splash of permissions poured into a generic “web access” switch.
reading: Panniantong/Agent-Reach
left on the table
- Plannotator’s human-question workflow: useful, but already covered in the 2026-07-06 Radar edition.
- edgeever: insufficient primary evidence of a distinct operating claim for this issue.
Related self.md routes
- Personal AI OS tools — the control-plane map for personal agents, receipts, memory, and tools
- AI coding assistants — compare coding workbenches by review surface, permissions, cost, logs, and escape hatches
- Sandboxing and security — give agents useful power without giving them the whole house