Skip to content

■ SIGNALS // RADAR SIGNAL

the agent passed its own test. that was the problem.

New papers test tool-using model refusals, a surrogate check for black-box agent calls, and coding agents that verify their own patches.

■ [!] ON THIS PAGE ▼

self.md radar — 2026-10-06

the agent passed its own test. that was the problem.

The failure is often not in the answer. It arrives one step later: after a model opens a tool, assigns a confidence score, or writes the test that approves its patch.

1. Tool use changed the refusal test

MLLMs Fail to Refuse when Using Tools Agentically puts multimodal models into a loop with visual tools such as zooming and tagging, then tests the safety boundary there. The authors report that refusal failures rose by as much as 68.7% relative to their no-tool setting across three safety benchmarks.

That number is a paper result, not a census of deployed agents. Still, it names a practical testing mistake: evaluating the model before it starts clicking around, then assuming the boundary survives the extra context and agency.

reading: MLLMs Fail to Refuse when Using Tools Agentically

2. The agent’s confidence is not much of a witness

Proxy Confidence proposes a spare model beside a black-box agent. It receives the same context, tool schema, and proposed call, then scores the call from its own probabilities. The authors report 0.825 AUROC on difficult coding tasks, where the acting model’s stated confidence sat near chance at 0.598.

The useful wrinkle is that wrong arguments and wrong actions are not the same failure. A method that can point at a suspect value is doing another job from one that says the whole call looks odd. Neither replaces review. It gives review a smaller pile.

reading: Proxy Confidence: Auditing Black-Box LLM Agents with a Surrogate’s Log-Probabilities

3. A patch needs evidence from outside the patch

Teaching Agents to Code Reliably starts with an awkward fact of coding-agent life: a test written by the agent for its own patch may happily approve the wrong repair. The paper trains search diversity and verification separately, including checks against a reverted tree; it reports verifier precision rising from 26.8% to 41.7% on its held-out evaluation.

The benchmark numbers will travel farther than the discipline behind them. For a live repository, the question is simpler: what would have to be true for this patch to be wrong, and did anything outside the patch get a chance to answer?

reading: Teaching Agents to Code Reliably

left on the table

  • omnigent-ai/omnigent: fresh repository, not enough source material to make a public claim without filling in the gaps.
  • Khanmigo’s two-year school experiment: the intake path did not provide enough primary-study detail for this edition.
  • Personal AI OS tools — the control-plane map for personal agents, receipts, memory, and tools
  • Agent memory systems — what agents should remember, what belongs in logs, and how to avoid memory sludge