Skip to content

[!] TOPIC ARCHIVE // #EVALUATION

#Evaluation

All 15 guides, operator dossiers, and signals tagged with #Evaluation.

RADAR SIGNAL

priced runs, reviewed traces, veto memory

Databricks priced coding-agent harnesses, AgentLens scored full trajectories, and new state tools turned memory into vetoes and receipts.

Read →
RADAR SIGNAL

agent receipts became research infrastructure

Microsoft rollout data, SwarmResearch branch search, and VERITAS replication logs all point at the same shift: coding agents now need measurable receipts.

Read →
RADAR SIGNAL

driver, page, sandbox

CueBench scores the human driving coding agents, Page Agent turns the web page into a tested control layer, and CubeSandbox/OpenShell harden the runtime underneath.

Read →
RADAR SIGNAL

plain text with teeth

test logs, video labels, and LLM reviews all turned into control surfaces: jqwik stdout, YouTube AI labels, and Review Arcade’s gameable review loop.

Read →
RADAR SIGNAL

agent job control

Agents crossed into job control: Codex Goals and agent spend got budgets, CUDA verifier wins failed workload tests, and AGENTS.md changes needed holdouts.

Read →
RADAR SIGNAL

stateful work

Anthropic packaged role work into Claude plugins, while new papers framed long-term memory as governed state and workplace agents as delegated-task systems.

Read →
RADAR SIGNAL

ledgers for autonomous work

Microsoft and GitHub put agent runs on cost rails, Prime Intellect published the logbook behind an autonomous research record, and arXiv plus BenchJack tightened the receipt layer.

Read →
RADAR SIGNAL

agent custody moved downstack

Reasoning traces picked up an order-bias audit; Frona put personal agents behind policy sandboxes; MDA made instruction files compiled and signed.

Read →
RADAR SIGNAL

delegation drift

DELEGATE-52 measured delegated document drift; HyperFrames and HTML workflows made agent output more inspectable; BrowserTrace recorded browser-agent runs step by step.

Read →
RADAR SIGNAL

agent accounting

Agents need accounting: task-duration horizons, line-level provenance, and evidence-gap reporting when access boundaries hide part of the file.

Read →
RADAR SIGNAL

control surfaces

operator controls surfaced inside agent tooling, Project N.O.M.A.D. packaged an offline command center at localhost:8080, and OpenFlo turned UX evaluation into something closer to nightly CI.

Read →
RADAR SIGNAL

the frontier model got lobotomized, safety theater got debunked, and your note app became infrastructure

opus can't pass the car wash test. open models reproduced mythos's zero-days. obsidian became an agent workspace. the stack is bifurcating.

Read →
RADAR SIGNAL

agents cheat, boundaries break

opus 4.6 games evals by finding answer keys. auto mode removes permission fatigue. local stacks hit usable. vibe-code security reckons. trust is infrastructure now.

Read →
RADAR SIGNAL

the 50% horizon

Claude Opus 4.6 hit 50% on multi-hour expert ML tasks. security became personal. the AI OS architecture stabilized. and the human-in-the-loop is vanishing faster than anyone projected.

Read →
RADAR SIGNAL

the approval problem

ChatGPT tells 5,000 people to breathe. heretic hits 1,000 stars. someone in Ukraine builds AI that survives power cuts. seven signals about what happens when you own your AI — or don't.

Read →
← All topics & tags