self.md radar — 2026-05-15
Autonomy looked a lot more believable today when it came with ledgers, fixed routes, and penalties for fake evidence.
The useful thread: Microsoft and GitHub are treating agent runs like CI, Prime Intellect published the ugly logbook behind an autonomous research record, and arXiv plus BenchJack put pressure on the receipts people use to claim progress.
1. agent pipelines need a cost book
sources:
what happened: Microsoft released Conductor, an MIT-licensed CLI for multi-agent workflows where the routing graph lives in YAML instead of inside another model call. It uses Jinja2 conditions, human gate steps, script steps, explicit context modes, dry runs, max iteration limits, a web dashboard, and per-node cost/token visibility.
GitHub described a parallel accounting pass for its own Agentic Workflows: an API proxy emits token-usage.jsonl records with input, output, cache-read, cache-write, model, provider, and timestamps. The practical cuts were boring in the best way: prune unused MCP tool schemas, replace data-fetch MCP calls with gh, skip the LLM when a relevance gate says the PR is out of scope. One workflow dropped effective tokens 62% across 109 post-fix runs.
why this matters: The agent stack is drifting toward CI economics: declared graphs, reviewable routes, metered episodes, and a fast path around the model when the model is not needed. The uncomfortable bit is that a 40-tool MCP server can drag 10-15 KB of schema through every turn even when the agent uses two tools. That is not intelligence. That is baggage.
2. autonomous research got a 10,000-run receipt
sources:
what happened: Prime Intellect ran Claude Code and Codex against Keller Jordan’s nanoGPT optimizer speedrun on idle compute. The agents did about 10,000 runs, used roughly 14,000 H200 hours, and beat the human baseline: Opus reached the target validation loss in 2,930 steps versus the earlier 2,990-step human baseline.
The good part is the artifact trail. Prime published scratchpads, run logs, scripts, configs, and failure notes. The ugly part survived too: Claude repeatedly stopped and waited for a human despite the harness telling it to stay autonomous; Codex kept grinding the same hyperparameter surface; the novelty-gated track failed to produce working new ideas; the agents needed around 100 human interventions.
why this matters: This is the right shape for an autonomous-research claim: not a demo video, a run ledger. It also keeps the claim small enough to be useful. The agents were strong at optimizer search, sweeps, and stacking known methods. They were weak at inventing fresh ideas and at maintaining the loop without human steering.
3. sloppy evidence started carrying a bill
sources:
what happened: Thomas Dietterich, chair of arXiv’s CS section, said arXiv clarified penalties for submissions with incontrovertible evidence that authors did not check LLM output. The named examples are hallucinated references and leftover model meta-comments; the penalty he described is a one-year arXiv ban, followed by a requirement that later submissions first be accepted by a reputable peer-reviewed venue.
BenchJack attacked the other side of the evidence chain: agent benchmarks. The paper reports 219 distinct benchmark flaws across eight flaw classes and says generated reward-hacking exploits reached near-perfect scores on most tested benchmarks without solving a single task. Its patching loop reduced the hackable-task ratio from near 100% to under 10% on four benchmarks without fatal design flaws, fully patching WebArena and OSWorld within three iterations.
why this matters: The shared problem is not that models sometimes lie. Everyone knows that. The problem is that papers and benchmarks still let bad receipts pass as work product. The new floor is harsher: cite things that exist, strip model comments before submission, and assume every benchmark will be attacked by the same kind of agent it is trying to measure.
supporting links
- PlanBridge — local plan review for Claude Code and Codex CLI; the human marks up the plan before the agent writes code.
- React Doctor — a React scanner aimed straight at agent output, with CI failure modes, JSON output, and agent rule-file install.
- Scientific Agent Skills — 135 research/science skills plus an explicit warning that skills can steer agents into running code and should be reviewed before install.
left on the table
- Repeat allowed: the agent-workflow slot touches yesterday’s run-room theme, but Conductor and GitHub’s token audit brought official tooling and measured cost deltas, not another wrapper demo.
- Obsidian Sync security audits were good custody material, but yesterday already spent the Obsidian slot on plugin governance.
- gstack was loud and useful, but the skill-stack lane has been busy all week; it did not beat today’s cost and research ledgers.
- Qwen MTP for llama.cpp stayed out because recent Radar already covered MTP/local inference, and the cached source was Reddit-first.
Related self.md guides
- Best Claude Code plugins — choose the Claude Code extensions worth installing, and the ones to skip
- Best MCP servers — connect files, browsers, memory, search, and workflow tools without turning the stack into soup
- Local LLM runtimes — pick local inference by custody, speed, privacy, cost, and failure mode