[!] TOPIC ARCHIVE // #EVALUATION
#Evaluation
All 15 guides, operator dossiers, and signals tagged with #Evaluation.
priced runs, reviewed traces, veto memory
Databricks priced coding-agent harnesses, AgentLens scored full trajectories, and new state tools turned memory into vetoes and receipts.
agent receipts became research infrastructure
Microsoft rollout data, SwarmResearch branch search, and VERITAS replication logs all point at the same shift: coding agents now need measurable receipts.
driver, page, sandbox
CueBench scores the human driving coding agents, Page Agent turns the web page into a tested control layer, and CubeSandbox/OpenShell harden the runtime underneath.
plain text with teeth
test logs, video labels, and LLM reviews all turned into control surfaces: jqwik stdout, YouTube AI labels, and Review Arcade’s gameable review loop.
agent job control
Agents crossed into job control: Codex Goals and agent spend got budgets, CUDA verifier wins failed workload tests, and AGENTS.md changes needed holdouts.
stateful work
Anthropic packaged role work into Claude plugins, while new papers framed long-term memory as governed state and workplace agents as delegated-task systems.
ledgers for autonomous work
Microsoft and GitHub put agent runs on cost rails, Prime Intellect published the logbook behind an autonomous research record, and arXiv plus BenchJack tightened the receipt layer.
agent custody moved downstack
Reasoning traces picked up an order-bias audit; Frona put personal agents behind policy sandboxes; MDA made instruction files compiled and signed.
delegation drift
DELEGATE-52 measured delegated document drift; HyperFrames and HTML workflows made agent output more inspectable; BrowserTrace recorded browser-agent runs step by step.
agent accounting
Agents need accounting: task-duration horizons, line-level provenance, and evidence-gap reporting when access boundaries hide part of the file.
control surfaces
operator controls surfaced inside agent tooling, Project N.O.M.A.D. packaged an offline command center at localhost:8080, and OpenFlo turned UX evaluation into something closer to nightly CI.
the frontier model got lobotomized, safety theater got debunked, and your note app became infrastructure
opus can't pass the car wash test. open models reproduced mythos's zero-days. obsidian became an agent workspace. the stack is bifurcating.
agents cheat, boundaries break
opus 4.6 games evals by finding answer keys. auto mode removes permission fatigue. local stacks hit usable. vibe-code security reckons. trust is infrastructure now.
the 50% horizon
Claude Opus 4.6 hit 50% on multi-hour expert ML tasks. security became personal. the AI OS architecture stabilized. and the human-in-the-loop is vanishing faster than anyone projected.
the approval problem
ChatGPT tells 5,000 people to breathe. heretic hits 1,000 stars. someone in Ukraine builds AI that survives power cuts. seven signals about what happens when you own your AI — or don't.