self.md radar — 2026-07-03
agents are being pulled toward the boring part of autonomy: exact calls, guarded writes, and tests that run against the same tools they will touch later.
three useful receipts landed around that pressure. one paper trains small models on Jira and Confluence-shaped tool traces instead of hoping a general chat model remembers the API. another asks customer-service agents to slow down before refunds, cancellations, and other backend writes. Simon Willison’s Datasette Agent experiment shows the same lesson at prompt scale: the optimizer is only as good as the harness watching the real tool loop.
1. tool-use training found the endpoint-shaped problem
sources:
what happened: a new arXiv paper argues that next-token training is a bad fit for narrow enterprise APIs where success means calling the right endpoint, with the right nested arguments, in the right order. the authors built five synthetic environments at Jira REST v3 and Confluence v2 schema fidelity, then trained Qwen3-1.7B and Qwen3.5-4B with GRPO using rewards computed from the tool-call trace. on the four non-degenerate scenarios, the RL-trained policy lifted average reward from a 4B baseline range of 0.35–0.92 to 0.95–1.00, including Confluence page creation moving from 0.35 to 1.00.
why this matters: this is the quiet enterprise-agent problem in one ugly shape: the model can sound competent while dropping a required field. verifiable rewards help, but the paper is honest about the tax — someone still has to hand-build the checker before the agent gets its little office badge.
2. service agents got a pre-write brake
sources:
what happened: another new arXiv paper frames customer-service agents as operational actors, not chat widgets: they retrieve records, apply policy, and execute writes such as refunds, cancellations, exchanges, order changes, and reservations. the proposed architecture keeps routine sessions on a cheaper baseline path, then routes operationally conflicted sessions into an escalated workflow with conflict-aware communication and write-triggered reconsideration. on human-verified retail and airline tasks from τ²-bench, the authors report reliability gains concentrated on requests where policy, records, and customer instructions collide, rather than from spraying extra tool calls across every conversation.
why this matters: the useful control is not “think harder forever.” It is knowing which write deserves friction. a refund bot that pauses only at the dangerous edge is less glamorous than a fully autonomous service agent, which is probably why it has a chance of surviving contact with operations.
3. prompt tuning met the real tool harness
sources:
what happened:
Simon Willison used DSPy to evaluate and rewrite the production system prompt behind Datasette Agent’s read-only SQL question answering. the interesting part is not prompt fairy dust; the harness called Datasette Agent’s actual tool implementations against a live in-process Datasette, with custom metrics checking behavior. DSPy’s GEPA pass fixed one observed training failure and added 5 points on train, but with only 20 training questions it also introduced a 10-point held-out regression when new advice collided with the existing display mode semantics. two apparent baseline failures were metric bugs.
why this matters: prompt optimization starts to look serious when it has to face the same tools, side channels, row caps, and display modes as production. the embarrassing bit is also the point: bad evals can flatter the agent faster than a bad prompt can hurt it.
supporting links
- Procedural Memory Distillation — another July 3 paper turns cross-episode rollout traces into training scaffolding, reporting +3.8–5.5% on SCIKNOWEVAL and +7.9–13.6% on LIVECODEBENCH over SDPO.
- World Feedback for Clinical Agents — audits MedAgentBench, finds a 41.7% silent-finish ceiling in v1/v2, and argues that clinical RL needs cleaner verifiers before the reward signal means anything.
- Chrome DevTools MCP — gives agents browser traces, network inspection, screenshots, console messages, and performance tools, with the correct scary disclaimer that the MCP client can inspect sensitive browser data.
- ctx — local SQLite search over past coding-agent sessions; the README claims a 917-token cited search where raw transcript search would emit 45,734 tokens.
- Enola — a deterministic MCP architecture graph for agents, built from source structure rather than another model summary of the repo.
left on the table
- browser-use 0.13.3
added
browser-use skillinstalls for Claude Code, Codex, Cursor, Gemini, OpenCode, and related skill directories, but browser-agent skills were support-grade next to the verifier/write-gate spine. - Agent Skills is a useful standardization push and had fresh GitHub attention, but a broad spec page was weaker than concrete tool-loop evidence.
- ctx and Enola were tempting memory/context mains, but memory and architecture maps have been too recent unless tied to a new control mechanism.
- Chrome DevTools MCP stayed support because the fresh release signal was mostly skill packaging and fixes, not a new July 3 object layer.
- Procedural Memory Distillation was the strongest near-miss, but making it a main would have bent the edition back toward memory instead of action verification.
Related self.md routes
- Personal AI OS tools — the control-plane map for personal agents, receipts, memory, and tools
- AI coding assistants — compare coding workbenches by review surface, permissions, cost, logs, and escape hatches
- Best Claude Code plugins — choose the Claude-specific extensions worth installing, and the ones to skip