Skip to content

■ SIGNALS // RADAR SIGNAL

May 25 Radar: serial reviewers, memory bills, and leaking sandboxes

Addy Osmani and Armin Ronacher put a human bottleneck under coding agents, Epoch and an energy paper put numbers under agent costs, and Canister shows why network allow-lists still leak secrets.

■ [!] ON THIS PAGE ▼

self.md radar — 2026-05-25

Today’s useful AI signal put the operator back inside the system diagram: agent work still queues at one human, infrastructure costs are drifting toward memory and retries, and a sandbox with an allowed endpoint can still leak the crown jewels.

Start with the human choke point: Addy Osmani wrote the clean name for it, and Armin Ronacher showed the issue-tracker version in Pi. Then the bill gets a denominator, first in HBM component spend and then in energy per successful goal. The security item is smaller but meaner: domain allow-lists stop destinations, not payloads.

1. Review stayed single-threaded

sources:

what happened: Addy Osmani published a May 24 post from a Google I/O panel point: spinning up more coding agents is cheap, but steering them, judging their output, and merging the changes still route through one human. His warning is not “use fewer agents”; it is that the scarce component in a multi-agent workflow is still the reviewer’s attention.

Armin Ronacher hit the same wall from the maintainer side while building Pi with Pi. Pi issue descriptions now double as prompt input for coding sessions, so the bad report has changed shape: “5% human and 95% clanker-generated,” plausible, confident, and wrong about root causes. Simon Willison pulled out the useful repair: issue reports should stay close to what the human actually observed — command, expectation, actual result, log.

why this matters: The next limit in coding agents is not how many workers you can launch; it is whether the human-facing input and review surface keeps its signal. If the prompt, issue, or diff review is already slop, parallelism just manufactures more cleanup.

2. The bill left the benchmark

sources:

what happened: Epoch AI’s updated chip-cost work says high-bandwidth memory rose from 52% to 63% of AI chip component spending between Q1 2024 and Q4 2025, averaged across chips from Nvidia, AMD, Google, and Amazon. In dollars, HBM spend across those designers grew from roughly $12B in 2024 to $32B in 2025, while total component spend rose from about $22B to $52B.

A new arXiv paper, “Energy per Successful Goal,” argues that invocation-level energy benchmarks break down for agents because a user goal can include tool calls, retries, and failure recovery. Its A-LEMS framework reports agentic workflows consuming 4.33× higher mean energy per successful goal than linear baselines across five reasoning and three tool-augmented task families — 888.1 J versus 205.3 J — while some tool-augmented cases invert below 1.0×.

why this matters: Agent economics are getting measured at the wrong seam if the denominator is still “one call” or “one answer.” The hard costs are piling up around memory supply, failed attempts, orchestration overhead, and completed work; the operator who tracks those beats the one who only watches leaderboard rank.

3. Allowed pipes still leak

sources:

what happened: André Graf published a Canister note on the blind spot in sandbox network allow-lists. The examples are blunt: an npm install allowed to reach its registry can leak credentials through DNS subdomains, and a build script can post a base64-encoded SSH key to an approved analytics endpoint. Nothing needs to “break out” if the secret rides through a destination the policy already permits.

The proposed repair is an L7 egress proxy with data-loss prevention, not just a domain filter. The scanner inspects DNS labels, headers, URIs, request and response bodies, and decoding chains such as base64, hex, percent-encoding, JSON escapes, HTML entities, and compressed payloads before traffic leaves.

why this matters: Agent sandboxes need to care about data flow, not just network shape. Once a tool can read files and call even one approved service, the safety question becomes what bytes crossed the boundary and whether anyone looked before they left.

  • Screenpipe app v2.4.271 — adds daemon-free screenpipe search for local history, then v2.4.273 removes Intercom live chat from desktop settings; personal-context tools are still sanding down the privacy edges in public.
  • DeepSeek-Reasonix — a DeepSeek-only terminal coding agent designed around prefix-cache stability, with a README case study claiming 435M input tokens, 99.82% cache hit, and about $12 for a day’s workload.
  • Context Drop — a tiny clipboard/file handoff tool for remote agent work, including an installable Agent Skills-compatible guide for Claude Code, Codex, Pi, and OpenCode.
  • NoteCast — local note engine that summarizes, extracts keywords, embeds, and organizes notes into an evolving graph; promising, but the README still says the code is not stable.

left on the table

  • The repeat pressure on Epoch’s HBM cost data was accepted because May 23 only carried a supporting memory-squeeze link; this was a fresh component-share data drop with dollar estimates.
  • Claude small-business skills stayed out because the Anthropic plugin/skill wave has been too hot this week and the cache only surfaced a Reddit report.
  • Chrome DevTools MCP did not repeat because it was yesterday’s main signal.
  • Datasette Agent 0.1a4 was fresh, but SQLite agent-addressability already carried the May 22 edition.