Skip to content

■ SIGNALS // RADAR SIGNAL

agent surfaces need boundaries

Open CTFs are losing their human scoreboard, html-anything turns local agents into HTML production workers, and Zerostack compresses the coding-agent harness into an auditable Rust shell.

■ [!] ON THIS PAGE ▼

self.md radar — 2026-05-17

The agent pressure showed up at the surface today: public scoreboards, publishable HTML, and the coding shell itself all had to admit the model is no longer a sidecar.

Open CTFs are losing the clean human ladder they used to provide. html-anything turns local coding agents into HTML production workers. Zerostack compresses the coding-agent harness into a small Rust binary with prompts, permissions, sessions, MCP, and optional sandboxing.

1. open CTFs hit the agent wall

sources:

what happened: Kabir Acharya argued that frontier models have broken the open online CTF format. His claim is not that AI can help with security work; CTF players have always used tools. The problem is that the model can now do enough of the reasoning, solve writing, and flag extraction that the public scoreboard stops measuring human skill cleanly.

The concrete shift is orchestration. Acharya says medium challenges became one-shottable earlier, then Claude Code-style CLIs made it trivial to wire a CTFd API, spawn model instances per challenge, and let agents burn down the easy and medium work while humans only touch what remains. Hacker News put the essay near the top of the front page with 371 points and 368 comments in the radar cache.

why this matters: This is the eval problem in hobbyist clothing. Once the public task can be copied into an agent loop, the scoreboard starts measuring budget, automation, and willingness to use frontier models. The old ladder from beginner struggle to expert status gets noisy. The same pressure hits any public benchmark that claims to measure human or model skill while its tasks remain fully open.

2. HTML became the agent output surface

sources:

what happened: nexu-io/html-anything crossed 2,500 GitHub stars in under a week. The pitch is blunt: “Markdown is the draft. HTML is what humans read. Your local agent writes it.” The app auto-detects eight coding-agent CLIs, including Claude Code, Cursor Agent, Codex, Gemini CLI, GitHub Copilot CLI, OpenCode, Qwen Coder, and Aider, then uses 75 skill templates across nine surface modes such as magazine articles, decks, posters, Xiaohongshu cards, tweet cards, prototypes, data reports, and Hyperframes.

The repo is not just another markdown renderer. It accepts Markdown, CSV, TSV, JSON, SQL, and plain text; streams the agent output into a sandboxed iframe; exports standalone HTML or high-DPI PNG; and has one-click paths for WeChat, X, Zhihu, and other publishing surfaces.

why this matters: The interesting move is away from chat as the artifact. A local agent is becoming a production worker that writes the final surface directly, with the editor providing skills, sandboxing, export targets, and enough design constraint to keep the output from looking like pasteurized slop. HTML is no longer just where the app runs. It is the handoff format.

3. coding agents got small enough to audit

sources:

what happened: zerostack, a Unix-inspired coding agent written in Rust, hit Hacker News and shipped as a crates.io package on May 16. The crate metadata describes it as a minimal coding agent optimized for memory footprint and performance; the README claims roughly 7k lines of code, an 8.9 MB binary, about 8 MB RAM on an empty session, and about 12 MB while working.

The feature list is the familiar modern harness, just compressed: OpenRouter, OpenAI, Anthropic, Gemini, Ollama, and custom providers; standard coding-agent tools; four permission modes; session save/load/resume with compaction; runtime prompt modes for code, plan, review, debug, ask, brainstorm, and security review; optional MCP; Exa-backed web tools; worktree movement; and optional bubblewrap sandboxing for shell commands.

why this matters: The coding-agent race is not only bigger IDE integrations. It is also smaller, inspectable harnesses where the control plane is visible: prompts, tools, permissions, context files, sessions, and sandbox policy. That matters because the risk has shifted from “will the model write code?” to “who owns the loop that lets it touch the machine?”

  • $δ$-mem — an online memory mechanism with an 8×8 state matrix, low-rank attention corrections, and reported gains on MemoryAgentBench and LoCoMo without extending the full context window.
  • llama.cpp MTP support — the multi-token prediction PR merged on May 16; the author reports roughly 75% steady-state acceptance with three draft tokens and more than 2× speedup over baseline in tested Qwen3.6 setups.
  • Semantic Notes Vault MCP — an Obsidian community-plugin post says the notes-vault MCP plugin finally made it through review after waiting since July 2025.

left on the table

  • The Bitwarden / Vaultwarden worry was a real custody story, but it leaned more business-model than agent-surface today.
  • MTP in llama.cpp could have taken a main slot, but yesterday already spent a main signal on local decoding speed via Orthrus. It belongs in the source trail, not the headline.
  • $δ$-mem was strong memory infrastructure, but the edition already had enough control-plane material and recent Radar covered memory as preflight.
  • OpenAI’s Malta partnership was broad adoption news, not a sharper operator signal than CTFs, HTML artifacts, or coding-agent shells.
  • Personal AI OS tools — the control-plane map for personal agents, receipts, memory, and tools
  • Sandboxing and security — give agents useful power without giving them the whole house
  • Best MCP servers — connect files, browsers, memory, search, and workflow tools without turning the stack into soup