self.md radar — 2026-07-07
coding agents had to leave the demo room and produce receipts: adoption curves, branch-level search, and replication logs.
one paper puts a number on Microsoft’s early CLI-agent rollout. another splits research agents across branches instead of letting one long context tunnel itself into a rut. a third uses coding agents to rerun papers and score claims against evidence, which is either beautiful or a very expensive way to rediscover peer review.
1. Microsoft put a rollout curve under Claude Code and Copilot CLI
sources:
what happened: Microsoft researchers studied an early-2026 rollout of Claude Code and GitHub Copilot CLI across tens of thousands of engineers. the paper says first use spread mostly through social networks, retention tracked coding activity more than demographics, and adopters merged roughly 24% more pull requests than the counterfactual estimate.
it also says token spend can reach millions of dollars a year at organizational scale, which is the sentence every procurement team will underline twice.
why this matters: the useful part is not “agents make engineers faster.” it is that CLI-agent rollout starts to look like a workplace contagion problem with a finance department attached: who sees peers using it, who keeps using it, and whether the extra PRs are worth the meter running in the background.
2. SwarmResearch treated parallel agents as search, not staffing
sources:
what happened: SwarmResearch argues that long-running coding agents can get stuck because one context and one editable program state make the search path too narrow. its harness uses a Shepherd Agent with global context to steer Search Agents working in separate git branches.
on open-ended optimization tasks, the authors report better or comparable solutions on 13 of 15 tasks versus other LLM-guided evolution and multi-agent baselines. the key detail is not just more agents. it is adaptive parallelism at different search depths.
why this matters: parallel coding agents are easy to turn into a very loud todo list. this version treats branching as epistemology: keep several possible worlds alive long enough for one of them to stop being stupid.
3. VERITAS made replication a CLI-agent job
sources:
what happened: VERITAS is a domain-agnostic replication framework built around CLI coding agents. given a paper, a code repo, or both, it extracts claims, runs the method while fixing breakage, judges each claim against experiment evidence, and returns a Replication Score, a severity-rated fix log, and a patched codebase.
its evaluation spans 65 papers across computer science, social science, medicine, and astrophysics. the authors say VERITAS beats two Claude Code baselines on CORE-Bench and ReplicationBench under the same model and host setup.
why this matters: this is the agent receipt pattern pushed into science: not “the paper seems plausible,” but here are the claims, here is the broken code, here are the patches, here is the score. the uncomfortable bit is that replication stops being a ceremonial virtue and starts leaving build artifacts behind.
supporting links
- Groundtruth — local Claude Code plugin that checks a “done” claim against the actual git diff, with no model call in the verdict path.
- Agentic OS — governance files plus hooks for secrets, tests, reviews, and skipped phases across Claude Code, Codex, Cursor, Copilot, and other markdown-reading agents.
- Schliff — deterministic scoring for AGENTS.md, SKILL.md, CLAUDE.md, and similar instruction files; basically lint for the text that quietly steers the agent.
- Stet — private evals for coding agents on real repo work, including model, harness, instruction, and reasoning-level changes.
left on the table
- the coding-agent rollout/cost theme came back only because the Microsoft paper is a fresh first-party data drop, not another wrapper launch.
- addyosmani/agent-skills stayed out because it was already in the seen ledger and today had stronger first-party measurement.
- ByteAsk fit the verification theme, especially for C/C++ with compilers and sanitizers, but it was a product page without enough outside proof for a main slot.
- PAI was the more personal self.md-shaped item, but the source was still mostly launch-page promise.
- AgentPost had a clean agent-mailbox primitive; it became a supporting mental note, not a full signal.
Related self.md routes
- Personal AI OS tools — the control-plane map for personal agents, receipts, memory, and tools
- AI coding assistants — compare coding workbenches by review surface, permissions, cost, logs, and escape hatches
- Best Claude Code plugins — choose the Claude-specific extensions worth installing, and the ones to skip