self.md radar — 2026-07-05
agentic coding stopped being a clever editor trick and started looking like an operating budget with merge policy, tool failure modes, and local data boundaries attached.
three useful reads today: one team published the ugly economics of agent swarms, one tool-builder caught frontier models breaking a stricter edit schema, and two local-first projects pushed assistants closer to the stuff people actually live in.
1. the coding agent became a line item
sources:
what happened: Lovable engineer Alexander Lebedev wrote up his agentic-coding spend: roughly $600/month before joining, about $25k in May, and around $85k since January. the process changed with the bill. a 20–30 PR week became a 150+ PR week, with one human steering 6–7 agents, a task-writing agent above them, and large changes moving as 10-PR stacks instead of single heroic branches.
human review did not vanish; it got moved. the post describes RFCs, ADRs, whiteboard sessions, and a markdown policy file that classifies PRs by size, risk, ownership, infra/auth surface, and production impact before GitHub rules decide whether the change gets fast AI review, slower AI review, or a human in the chair.
why this matters: that is the honest shape of scaled agent work: not “AI writes code,” but token spend, review lanes, and a boring policy file deciding when the human gets paged. the self.md angle is painfully practical: if agents are going to ship more than toys, the personal operating system needs accounting and escalation, not vibes in a chat tab.
2. better models broke the edit tool
sources:
what happened:
Armin Ronacher hit a Pi bug where newer Claude models called an edit tool with extra invented fields inside the nested edits[] array. the edit intent was usually right, but the payload failed the schema, so Pi rejected it and asked the model to try again. the nasty part: he saw the regression in Opus 4.8 and Sonnet 5, not in older siblings.
his point is simple and annoying: tool calls are still text wearing a little ceremonial robe. the harness validates the shape after the model emits it, and stricter schemas can turn a smarter model into a worse tool user. HIC Mouse is the commercial counter-move in miniature: coordinate edits, staged rollback, and “tool-response engineering” instead of trusting one string-replace spell to hold a repo together.
why this matters: the frontier-model story keeps pretending capability is monotonic. tool work says otherwise: a better model can be worse at the exact interface your system depends on, and the interface needs its own engineering budget.
3. local context got teeth
sources:
what happened: Local MCP is pitching a Mac bridge that lets ChatGPT and Claude on the web touch local apps through 183 tools: Mail, Calendar, iMessage, Teams, Slack, WhatsApp, Signal, Finder, Notes, Office, Drive, and more. the promise is no API tokens for local personal data and no cloud middleman for the app contents. that is either exactly what personal AI needs or the beginning of a very spicy permissions dialog, depending on how awake the operator is.
OpenBiliClaw comes from the other side of the same wall. it runs locally, stores the user profile in SQLite, and uses feedback plus browsing behavior to send an agent across Bilibili, Xiaohongshu, Douyin, YouTube, X, Zhihu, Reddit, and the web to find content on the user’s behalf. it is not just “recommend me videos”; it is a tiny local taste bureaucracy with a browser extension and a psychological profile.
why this matters: personal AI only gets useful when it can see the messy private layer: messages, files, calendars, taste, drift. the cost is that “local-first” stops being a privacy sticker and becomes a permissions architecture.
supporting links
- Greplica memory benchmark — repo memory claimed 43% lower cost, 49% fewer tokens, 36% fewer tool calls, and 26% less planning time on 10 SWE-chat tasks.
- Mouse precision editing — worth opening beside Armin’s piece because it treats file editing as a dedicated agent interface, not an afterthought.
- Local MCP HN thread — the source discussion around local app access is probably more useful than the landing-page perfume.
- OpenBiliClaw changelog — the live edge is mundane but revealing: broken vector models, deep links, restart recovery, translated settings corrupting provider values.
left on the table
- Alibaba page-agent stayed out because it was already used yesterday; same URL, same browser-agent wave.
- openai/codex-plugin-cc also stayed out after yesterday’s support slot. useful tool, bad repeat.
- Chrome DevTools MCP and Agent Skills were still too recent from July 3 to drag back into the main edition.
- Immich and Romm had traffic, but not enough self.md-specific operator consequence for a main signal.
- Odin/Wikipedia engagement farming was culturally interesting and strategically off-thesis. not every shiny annoyance gets a badge.
Related self.md routes
- Personal AI OS tools — the control-plane map for personal agents, receipts, memory, and tools
- AI coding assistants — compare coding workbenches by review surface, permissions, cost, logs, and escape hatches
- Best Claude Code plugins — choose the Claude-specific extensions worth installing, and the ones to skip