self.md radar — 2026-06-30
agent work is getting less mystical and more inspectable: traces for what actually happens, permission profiles for what the agent may touch, and local orchestrators for the ugly stretch between prompt and PR.
that gives today a nice unglamorous shape. UW measured the real loop and found shell-heavy, long-tailed agent work. OpenAI pointed Codex users toward beta filesystem and network permission profiles instead of wishful ignore-file folklore. DoorDash’s agentico release treats the coding agent as a workflow with artifacts, gates, costs, and a draft PR button.
1. the trace finally caught the agent loop
sources:
what happened: UW’s SyFI Lab published TraceLab, a public trace-and-analysis pipeline for coding-agent workloads, built around real Claude Code, Codex, and Gemini CLI usage instead of neat benchmark tasks. the post says each user request averaged 8.8 self-directed LLM/tool cycles and 10.8 tool calls before the final answer. 88% of LLM rounds responded to a tool result rather than a human, and 76% of 433K tool calls were shell or command executions. latency was ugly in the way real work is ugly: median end-to-end time around 38 seconds, mean around 4 minutes, and p99 nearly 44 minutes.
why this matters: coding agents are not just token streams with nicer autocomplete. if the loop is mostly tools, shells, waits, retries, and human gaps, the personal AI OS needs scheduling, tracing, cache policy, and cost accounting that match the mess.
2. Codex got pointed at a permission profile, not a prayer
sources:
what happened:
OpenAI closed the long-running Codex request for a way to keep agents away from sensitive files after pointing to its beta permission-profile system. the original issue asked for deterministic repo-local and global rules around .env, .env.*, *.pem, id_*, .aws/**, and .ssh/**, instead of relying on project docs and vibes. the docs now describe default_permissions and [permissions] profiles for filesystem and network access, including rules like workspace write access, .devcontainer read-only access, **/*.env denied, and per-domain network allow/deny. OpenAI is careful to call it beta and separate it from older sandbox settings, which is exactly the kind of boring caveat this feature needs.
why this matters: ignore files are etiquette. permission profiles are a control surface. the useful question is no longer “did the agent promise not to look?” but “what did the runtime make impossible?”
3. the coding agent grew a project manager-shaped shell
sources:
what happened:
DoorDash’s Agentic Orchestrator frames agentico as a local CLI that turns a feature request into research, planning, implementation, review, and a pull request from one terminal. the README is blunt about the failure mode: asking a model to edit files is not the hard part; getting from a vague request to a reviewable PR without dropping context or letting a bad plan inflate the diff is. it persists inquiry, research, design, roadmap, and phase-plan artifacts, runs separate worktrees and sessions, and treats Claude, Codex, and OpenCode as phase-selectable providers. the 2026-06-29 v0.147.0 release adds draft PR publishing, separate Inquire and Research models, a planning model for design, product-framed final review, and more accurate cost accounting for review sessions, Codex cached inputs, and OpenCode child sessions.
why this matters: the operator object is shifting from “chat with a coding model” to “supervise a state machine.” that is less romantic, which is good. romance is how you get a 4,000-line diff and no idea who approved the plan.
supporting links
- Hamel Husain on evals as product smell — useful pressure on the same point: if evals are hard, the product surface is probably too vague.
- CI/Lock — SLSA-aligned in-toto attestations for CI, scans, and compliance; not agent-specific, but a clean receipt layer for the work agents increasingly touch.
- Vercel Labs skills
—
npx skillsnow sits in the open-agent-skills lane, with support for Claude Code, Codex, Cursor, OpenCode, and a long tail of agents. - MCP Compressing Router
— turns many MCP servers into one router with
get_tool_schemaandinvoke_tool, claiming a 26K-token tool-description overhead can drop to about 900 tokens in its example. - screenpipe MCP v0.18.14 — one-click MCP OAuth for 10 connectors, plus enterprise fixes around managed settings and context-display leaks.
left on the table
- Ornith-1.0 had HN attention again, but yesterday already used the model/harness story as the lead.
- Katra and MateClaw stayed out because another memory-stack signal would be a repeat without a clean new object layer.
- ClawSec was conceptually on-theme, but the fresh activity was packaging/wiki-side, not a meaningful new control.
- Gusto’s Claude Code shipping story was vivid, but anecdote lost to trace data and live tooling.
- South Korea’s memory-and-robots spend was big infrastructure news, not a direct operator surface for this edition.
Related self.md routes
- Personal AI OS tools — the control-plane map for personal agents, receipts, memory, and tools
- AI coding assistants — compare coding workbenches by review surface, permissions, cost, logs, and escape hatches
- Best Claude Code plugins — choose the Claude-specific extensions worth installing, and the ones to skip