self.md radar — 2026-05-27
agent work got a little less romantic today. not worse. just more file-system, more receipts, more “show me the state change.”
Anthropic shipped role work as plugin folders, not vibes in a prompt doc. two new papers did the other half of the boring adult work: memory as governed mutation, job agents as delegated tasks with evidence rubrics. this is the useful version of the hype getting house keys and immediately discovering chores.
1. role work moved into folders
sources:
what happened:
Anthropic’s knowledge-work-plugins repo was active today and open-sources 11 role plugins for Claude Cowork, also installable in Claude Code. this is not another prompt gallery. each plugin is a little bundle of instructions and permissions: .claude-plugin/plugin.json, .mcp.json, commands/, and skills/, covering productivity, sales, support, product, marketing, legal, finance, data, enterprise search, bio research, and plugin management. the install path is explicit too: add the anthropics/knowledge-work-plugins marketplace, then install a role plugin like sales@knowledge-work-plugins.
why this matters: the interesting part is the control surface. prompts become installable, reviewable, forkable work packages — much closer to operating an AI workbench than teaching everyone in the company the same little spell.
2. memory needs edits, not just retrieval
sources:
what happened:
Is Agent Memory a Database? argues that long-running agents break the normal database metaphor. the authors say current memory systems mostly localize correctness around records, embeddings, or graph edges, then quietly inherit four bad habits: unregulated growth, missing semantic revision, capacity-driven forgetting, and read-only retrieval. their proposed abstraction, Governed Evolving Memory, moves correctness up to the trajectory of the memory state, with four state-level operators: ingestion, revision, forgetting, and retrieval.
why this matters: a memory system that only retrieves is a diary with a search bar. useful agent memory needs receipts for what changed, why it changed, what got forgotten, and which belief is still allowed to touch the steering wheel.
3. JobBench made delegation annoying again
sources:
what happened:
JobBench moves occupational-agent evaluation away from “can this model cosplay a job title?” and back toward the uglier unit: can it finish the thing a human would actually delegate? the benchmark covers 130 agentic tasks across 35 occupations, with each task packaged as a messy workspace of heterogeneous reference files. outputs are graded by fact-anchored rubrics averaging 35.6 binary criteria per task, and the strongest evaluated setup — Claude Opus 4.7 under Claude Code — reaches 45.9%.
why this matters: that number is a useful cold shower. the real test is not whether an agent sounds like an analyst, lawyer, or marketer; it is whether it can handle clutter, constraints, and evidence without leaving a small corpse for the human to clean up.
supporting links
- Claude Code daily-driver guide
— practical companion reading on
CLAUDE.md, skills, subagents, plugins, MCPs, and verification loops. - claude-external-memory — a small AGENTS/PLAN/PROGRESS/VERIFY pattern for repo memory; good concrete foil to the memory paper.
- DeepSWE — contamination-free long-horizon coding benchmark; same pressure as JobBench, just pointed at coding work.
- Vellum Assistant — personal assistant repo with memory, personality, and proactive reach-outs across macOS, Telegram, and Slack.
left on the table
- Manifest had the biggest GitHub score, but smart model routing is an older repo and the cost-control lane has been hit recently.
- Understand Anything stayed out because the code-graph/workbench wave already showed up in recent Radar history.
- cc-switch is useful wrapper consolidation for Claude Code, Codex, OpenCode, OpenClaw, Gemini CLI, and Hermes Agent, but not a new control shift.
- Binary and Ternary Bonsai Image is technically fun, especially the browser/WebGPU angle, but it belongs to the local media-model lane, not today’s agent-ops spine.
Related self.md routes
- Personal AI OS tools — the control-plane map for personal agents, receipts, memory, and tools
- AI coding assistants — compare coding workbenches by review surface, permissions, cost, logs, and escape hatches
- Best Claude Code plugins — choose the Claude-specific extensions worth installing, and the ones to skip