Skip to content

■ SIGNALS // RADAR SIGNAL

priced runs, reviewed traces, veto memory

Databricks priced coding-agent harnesses, AgentLens scored full trajectories, and new state tools turned memory into vetoes and receipts.

■ [!] ON THIS PAGE ▼

self.md radar — 2026-07-09

agent work got dragged out of the vibes layer today and into the parts teams can price, review, and write down.

the useful cluster is all plumbing: Databricks priced model-plus-harness choices on real code tasks, AgentLens scored whole trajectories instead of a green check, and new state tools tried to make agents remember decisions as vetoes and receipts rather than smoke in a drawer.

1. harness choice turned into a bill

sources:

what happened: Databricks published an internal coding-agent benchmark built from actual engineer tasks on its multi-million-line codebase, with tasks and solutions reviewed before scoring. the uncomfortable part was not just which model won: GLM 5.2 landed statistically tied with Opus 4.8 on quality while costing $1.28 per task against Opus at $1.94, and Sonnet 5 ended up costing more per task than Opus because it used 1.9x more tokens while scoring six points lower. a separate Harness Effect paper did a controlled swap across 22 locked tasks and six models, changing only the orchestration layer; its harness cut blended cost per task 41%, wall-clock time 44%, and tokens per task 38%, with quality roughly flat. Lenny’s crew then made the same word practical with a Sentry bug-triage harness wired into Linear, GitHub, Vercel, a terminal UI, permissions, and reusable artifacts.

why this matters: Model choice is now only one line in the receipt. the expensive part is how the system feeds context, exposes tools, asks for permission, and turns a run into something the rest of the team can use.

2. agent eval started reading the whole mess

sources:

what happened: AgentLens argues that pass/fail coding-agent benchmarks miss the part users actually suffer through: the trajectory. its benchmark records full headless-IDE interactions, then combines formal checks with LLM-written reviews and side-by-side comparisons. the open-source fold ships 16 Java workflow scenarios and judge dimensions that include end result, instruction compliance, pitfalls, pleasantness, and tool calls. the paper frames this as production-assessed evaluation: useful for ranking, but also for nightly regression checks when an agent gets worse in a way a single green test would hide.

why this matters: As agents become coworkers with terminals, the run itself becomes evidence. a failed task can still contain useful work, and a passing task can still be a little goblin that ignored instructions, wasted tool calls, or made the human babysit it.

3. memory tools learned the word no

sources:

what happened: RoBrain shipped as team memory for agents, but its sharp bit is the veto: decisions carry a structured rejected[] field, and its VetoBench claim is blunt — without decision memory, an agent re-proposes a rejected approach in up to 9 of 10 tasks; through the RoBrain pipeline, it reports 0 of 50 violations across five archived runs. Cruxible comes at the same scar from state instead of memory: typed entities, deterministic workflows, governed writes, evidence-backed proposals, human review where trust rules cannot decide, and receipts on accepted claims. skill-extractor mines Claude Code, Codex CLI, and other agent transcripts for repeated fixes, scores candidates by confidence and utility, redacts risk, then asks a human before writing a real SKILL.md.

why this matters: This is the more honest version of agent memory: not “remember everything,” but “remember what we already ruled out, who accepted a claim, and which correction deserves to become law.” recall is nice. vetoes are where the rework dies.

  • BrowserOS — an open-source Chromium fork with 53+ browser automation tools, an MCP server, scheduled tasks, and local-provider options; very relevant, but too close to this week’s browser-agent lane for a main slot.
  • Google Analytics MCP — Google Analytics data and admin APIs exposed as local MCP tools; useful proof that boring business systems are getting agent handles.
  • AI Readiness Inspector — checks llms.txt, robots.txt AI rules, structured data, meta tags, and /.well-known/agent.json; good operator footnote after yesterday’s crawler-access story.
  • Dreamer — mines recurring corrections from 20 coding harnesses and proposes context-file updates with evidence, which rhymes with the skill/state signal without carrying it.
  • Omnigent — a meta-harness for Claude Code, Codex, Cursor, OpenCode, Hermes, Pi, policies, sandboxes, and multi-agent supervision; watchlist, not today’s proof object.

left on the table

  • Memory in the Loop and TencentDB Agent Memory stayed out because yesterday already covered memory placement and loop-time latency.
  • agent-skills stayed out as an exact recent repeat from 2026-07-07; today’s skill item is about mining corrections from transcripts, not another static skill catalog.
  • jcode had harness fit, but the Databricks and Harness Effect sources carried stronger measured evidence.
  • Google’s vibe-coding SDLC whitepaper had big adoption numbers, including 85% regular coding-agent use and 51% daily use, but the piece was broader than today’s operator-grade thread.