Skip to content

■ SIGNALS // RADAR SIGNAL

local outputs, remote witnesses

Paint’s server-issued image identifier, Google’s extensively tested Rust rewrite, and a benchmark study where harness settings choose the model winner.

■ [!] ON THIS PAGE ▼

self.md radar — 2026-08-25

some of today’s supposedly local tools have a second address. Paint can run an image model on the machine while still asking a Microsoft service to moderate the prompt and issue a watermark identifier. Google’s Rust rewrite got through because it was tested against 30 million GIFs, not because Gemini wrote Rust. And a new benchmark paper shows that the harness can choose the winner before the model has done anything interesting.

1. Paint’s local image generator still asks Microsoft for a GUID

A reverse-engineering report by Xusheng traces the path through current Microsoft Paint and Photos on Copilot+ PCs. The local image-generation flow sends the prompt and style to a Microsoft moderation endpoint, receives a revised prompt plus IDs, then embeds the returned watermarkId into the pixels of the locally generated image. The report identifies a 16-byte GUID payload and shows that a failed invisible-watermark call turns the generation into an error rather than returning an unmarked file.

This is separate from Paint’s visible Copilot-logo setting. The researcher found a second path in Watermarker.dll, including a 144-bit payload written repeatedly into the image. The generated image can still carry C2PA metadata, which Microsoft documents for its AI image features . But C2PA metadata is only part of the story if the pixel-level mark persists after somebody strips metadata or exports another copy.

The report is independent research, not a Microsoft security advisory, so the sensible reading is narrow: this is what the analyst observed in the examined build, not proof of a universal Windows surveillance scheme. Still, “runs locally” has become a bad shorthand. A model file on the disk does not settle where the prompt went, who minted the identifier, or what remains inside the final object.

reading: the reverse-engineering report · Microsoft’s Paint AI documentation

2. Google’s Rust rewrite needed a test corpus the size of a small weather system

Google’s Security Blog describes an AI-assisted rewrite of giflib, a roughly 3,000-line C library that processes untrusted GIF data, into a Rust replacement compatible with existing C callers. Gemini helped create the initial translation, then the difficult part arrived: the FFI boundary was initially unsound, experts reviewed the remaining unsafe code, and differential-test failures were fed back into targeted fixes.

The resulting fork was checked against more than 30 million GIFs. A differential fuzzer ran for over six days and completed more than 200 million iterations without a logic deviation, according to Google; the company has also published the Rust implementation . That is the useful number in this story. The language conversion was fast. Making a replacement safe enough for production required an enormous witness set, a rollback plan, and service owners willing to trust the evidence.

There is no free migration hidden here. Google notes that a language-level fork creates ongoing maintenance work when upstream changes. But the post is a clean antidote to “AI rewrote it” as a finish line. In consequential code, generation is the cheap scene. Compatibility, observability, and the right to retreat are the actual job.

reading: Google’s engineering account · the published giflib-rs fork

3. a leaderboard score can be a property of the harness

A new arXiv paper holds model weights, questions, and greedy decoding fixed, then runs 12 open-weight instruction-tuned models across 3,679 benchmark items under 26 defensible harness configurations. Those configurations change option order, prompt wording, and the way answers are scored. The result is ugly for anyone treating a leaderboard as a neat ordering: the paper reports Gemma 4 31B ranging from 31% to 89%, and four of the 12 models taking first place under at least one configuration.

The authors say 95.7% of the average gap between adjacent models sits in configuration-fragile items. Their analysis and per-item records are released, so the claim can be inspected rather than merely admired. It remains one paper, using multiple-choice benchmarks and open-weight models; it does not prove that every evaluation is theatrical. It does show that a score without its harness is missing part of the instrument.

This matters for buying, deploying, or even casually comparing models. A benchmark result is not a property sticker on a box. It is a report from a particular arrangement of prompts, ordering, scoring rules, and code. If changing that arrangement chooses the winner, the arrangement belongs in the receipt.

reading: the paper and its methods · the experimental HTML version

more to read

  • Ambient Context — a new macOS menu-bar app writes focused-window text into daily Markdown files for an LLM, with local storage and pre-write redaction; promising, but its Accessibility permission is the whole proposition.
  • SELF — Farid Zakaria’s executable SQLite prototype makes symbols, dependencies, and stripping queryable; a lovely systems experiment rather than a replacement for ELF tomorrow morning.
  • AIREP — an experimental proposal for signed, hash-chained records of individual AI runtime decisions. The format is young, but it is at least trying to make “the system blocked this” into something a stranger can check.

left on the table

  • Ambient Context is fresh and unusually close to the personal-memory question, but an eight-star, same-day repository does not yet outrank three sources with a fuller public record.
  • AIREP has a reference implementation and conformance kit, but it is explicitly experimental and lands too close to last week’s run-log and receipt material to lead without adoption evidence.
  • the SQLite executable prototype is the day’s best weird object. It stays on the shelf because a prototype file format is not yet a change in how other people ship or maintain software.
  • Personal AI OS tools — the control-plane map for personal agents, receipts, memory, and tools
  • Local LLM runtimes — pick local inference by custody, speed, privacy, cost, and failure mode
  • Agent memory systems — what agents should remember, what belongs in logs, and how to avoid memory sludge