Skip to content

■ SIGNALS // RADAR SIGNAL

agent job control

Agents crossed into job control: Codex Goals and agent spend got budgets, CUDA verifier wins failed workload tests, and AGENTS.md changes needed holdouts.

■ [!] ON THIS PAGE ▼

self.md radar — 2026-05-28

the invoice finally caught up with the workbench.

agent work is acquiring the boring parts that make it real: budgets, workload checks, and change control. OpenAI turned Codex tasks into scoped goals, Simon Willison traced the agent-spend shock, doubleAI found a benchmark-fast CUDA kernel that breaks a small training run, and Stet’s AGENTS.md experiment improved the warmup set while failing the holdout.

1. the agent got a budget

sources:

what happened: Simon Willison argues OpenAI and Anthropic have crossed into product-market fit because internal agent use is starting to show up as real usage bills. his own last 30 days of Claude Code and Codex would have priced at $1,199.79 and $980.37 at API rates, while his subscriptions cost $200. OpenAI’s Goals guide puts the same shift into the product surface: a Codex goal is a scoped completion contract with a measurable outcome, a verification surface, constraints, /goal pause, /goal resume, /goal clear, and a budget stop. Claire Vo’s demo takes it out of docs and into long-running work: Sentry errors, Vercel API errors, Linear cleanup, and a 3,900-email inbox burn-down.

why this matters: a task that can run overnight and spend real money needs job controls, not just a nicer prompt box. the personal AI OS version of this is simple: every delegated job needs an owner, a budget, a stop condition, and receipts.

2. fast kernels need dirty workloads

sources:

what happened: doubleAI says WarpSpeed beat NVIDIA’s optimized PyTorch baselines on 90% of SOL-ExecBench’s 235 Blackwell kernels, running 2.24x faster on average after one day of search. then they took a verifier-passing embedding lookup plus RMSNorm backward kernel and put it inside a small transformer training run. on a Zipfian corpus with plain SGD, one kernel’s loss diverged. the bug was not dramatic: the fast version accumulated gradients directly in bf16, while the reference used an fp32 scratch buffer and cast once at the end. AdamW could mask the bias, which is exactly why this kind of bug wastes research time.

why this matters: an agent can optimize a kernel into a lab pass and a production lie. if the verifier does not model the workload, the failure gets misfiled as a bad idea, a weird dataset, or researcher incompetence.

3. AGENTS.md found its overfit problem

sources:

what happened: Stet had Codex iterate on its own AGENTS.md eight times against real historical repo tasks. the setup was small but honest: Codex with gpt-5.5 at medium reasoning, a gpt-5.4 grader, an n=5 training slice, and an n=10 clean holdout. the best candidate improved the training slice, then regressed on the holdout. footprint widened, tokens climbed, tool calls climbed, and code-review correctness fell, even while tests held flat.

why this matters: the instruction file is now a patch, not lore. if it ships without a holdout, nobody knows whether the agent got better or just learned the last five tasks.

  • Anthropic: How we contain Claude across products — gVisor containers, Claude Code approval fatigue, Cowork VMs, egress controls, and the ugly lesson that an allowed domain is also a capability grant.
  • VAEN — packages instructions, skills, and MCP config into portable .agent files without bundling secrets.
  • DeepSWE — another measurement track for frontier coding agents, useful companion pressure to Stet’s repo-specific eval story.
  • Agyn — an agent-definition-as-code paper with scalable execution and zero-trust access as the frame.

left on the table