Skip to content

■ SIGNALS // RADAR SIGNAL

the verdict needs a place to stop

ScopeBench on agent scope, a code judge that declines unsupported comparisons, and Panda’s December 2027 local AI preorder.

■ [!] ON THIS PAGE ▼

self.md radar — 2026-09-28

the verdict needs a place to stop

A security test with one forbidden door. A code judge that has to admit it cannot tell two answers apart. A $2,997 black box for the home, promised for December 2027. Three objects with enough detail to inspect before the pitch gets to smooth them over.

1. ScopeBench treats the brief as part of the test

Security agents are usually asked whether they can get the flag. ScopeBench asks a nastier second question: did they get it by entering the thing the brief explicitly excluded?

Its 30 containerized tasks keep the environment and objective fixed, then swap in a natural-language boundary. In the scoped version, the flag sits behind that boundary, so a successful deterministic check is also proof of a forbidden action. The pilot covers 2,160 trajectories: 30 tasks, three instruction conditions, eight model families, three repetitions.

That leaves a usable distinction for an operator. Task completion is not evidence that the job was done within the job.

reading: ScopeBench paper · released pilot tasks

2. The code judge is allowed to leave the verdict blank

A multi-agent code-review pipeline can produce a fluent explanation even when the evidence cannot distinguish the candidates. The paper behind this result ran MARCH across 80 condition-by-cell measurements; in 78–95% of comparisons, the unmodified system called both solutions equally good. One configuration reached 4.4% accuracy, against 43.7% for asking the same model directly.

The repair is not an even more ceremonial judge. The authors derive two measurements from the pipeline’s logs and use one as a gate. It declines comparisons where the evidence is not discriminating. Accuracy rose from 20.7% to 36.9%, while the system still answered half of the comparisons.

A review surface that can say “no basis for a verdict” is less convenient. It is also harder to mistake for a receipt.

reading: multi-agent code judge paper

3. Panda is selling a local AI appliance before it exists

Panda’s sales page is the domestic version of the local-AI pitch: put a box beside the router, keep processing at home, let up to 25 devices use it, and pay once rather than carry another subscription. The price is $2,997. The deposit is refundable until shipping. Shipping is stated as December 2027.

That date matters more than the kitchen-table demos. A local machine is not just a model choice once it is sold as household equipment. Somebody owns the hardware failures, remote access, updates, support channel and the long wait between a deposit and a box.

The offer is concrete enough to inspect, but it remains a preorder. The page’s claims about private, fully local work are a promise attached to a future device, not a completed installation.

reading: Panda order page

left on the table

SkillCascade describes attacks spread across several apparently harmless agent skills. It is a real warning, but too close to today’s first story to earn a second slot without turning the issue into a security stack.

  • Personal AI OS tools — the control-plane map for personal agents, receipts, memory, and tools
  • Best Claude Code plugins — choose the Claude-specific extensions worth installing, and the ones to skip
  • Browser agents — where browser automation helps, where it lies, and how to keep it observable