self.md radar — 2026-05-20
Google put Gemini 3.5 Flash straight into its agent surfaces while the best benchmark in the stack today said delegation can fail quietly behind a decent final score.
So this one is about the plumbing around agent work: a faster paid model, a routing test that looks past pass/fail, and a small local-model repo that treats bad tool calls as something to catch, not something to pray away.
1. Google shipped Flash into the agent lane
sources:
what happened:
Google released Gemini 3.5 Flash directly to general availability across the Gemini app, AI Mode in Search, Google Antigravity, the Gemini API in AI Studio, Android Studio, Gemini Enterprise Agent Platform, and Gemini Enterprise. The company says it beats Gemini 3.1 Pro on Terminal-Bench 2.1 (76.2%), GDPval-AA (1656 Elo), MCP Atlas (83.6%), and CharXiv Reasoning (84.2%), with 4x faster output tokens than other frontier models. Simon Willison caught the developer bill: gemini-3.5-flash, a 1,048,576-token input window, 65,536 max output tokens, $1.50/M input, $9/M output, and a price jump over older Flash models.
why this matters: Google wired the release into agent surfaces on day one. Faster output helps, but the price increase makes unattended runs a budget problem, not just a latency problem.
2. DecisionBench caught delegation hiding in the averages
sources:
what happened:
DecisionBench proposes a benchmark substrate for long-horizon agent delegation across GAIA, tau-bench, and BFCL multi-turn, with 11 peer models from 7 vendor families. The interface gives agents call_model and, optionally, read_profile, then measures quality, cost, latency, delegation rate, routing fidelity-at-k, vendor self-preference, and a counterfactual delegation ceiling. In a five-condition sweep over 23,375 task instances, mean end-task quality was statistically indistinguishable across awareness conditions, while routing fidelity@1 ranged from 7.5% to 29.5% and the perfect-delegation ceiling sat 15–31 percentage points above measured performance.
why this matters: An orchestrator can look fine on final score while calling the wrong peer. Agent products need routing logs, counterfactual baselines, and cost/latency receipts, or they are just grading the last mile and ignoring the handoff.
3. Forge put small-model reliability outside the model
sources:
what happened:
Forge is a Python reliability layer for self-hosted LLM tool-calling and multi-step workflows. Its README says a Ministral-3 8B Instruct Q8 setup on llama-server reaches 86.5% across a 26-scenario eval suite and 76% on the hardest tier when wrapped with rescue parsing, retry nudges, step enforcement, VRAM-aware budgets, and tiered compaction. It can run as a workflow runner, middleware inside another orchestrator, or an OpenAI-compatible proxy that injects a synthetic respond tool and strips it before the client sees the response.
why this matters: This is the sane local-agent bet: put discipline around the model instead of pretending an 8B checkpoint suddenly learned tool hygiene. The cheaper model becomes useful only when the surrounding loop is opinionated enough to say no.
supporting links
- Academic Research Skills — 14.5k-star Claude Code skill pack for research, writing, review, citation checks, and integrity gates.
- Cookiy user research skill — user interviews, transcript synthesis, surveys, and participant workflows exposed as an agent skill.
- 12-factor-agents — a production-agent checklist worth keeping beside any release-day speed claim.
left on the table
- Academic Research Skills stayed supporting because skills-as-packaging already had recent main coverage; Forge was the cleaner repo gift today.
- Trustworthy Agent Network , ReElicit , and Learn-by-Wire were all fresh, but DecisionBench had the sharpest operational metric.
- RTK was skipped as a main item because the repo was already covered in February and today did not add a hard product change.
- Karpathy-to-Anthropic chatter had heat, not enough operator consequence.
Related self.md routes
- Personal AI OS tools — the control-plane map for personal agents, receipts, memory, and tools
- AI coding assistants — compare coding workbenches by review surface, permissions, cost, logs, and escape hatches
- Best Claude Code plugins — choose the Claude-specific extensions worth installing, and the ones to skip