Skip to content

■ SIGNALS // RADAR SIGNAL

the cost of witnesses

OpenAI’s reported monitoring overhead, Mojo’s newly open compiler, and a clinical-trial DAG: three places where the process stops being decorative.

■ [!] ON THIS PAGE ▼

self.md radar — 2026-08-19

two pieces of infrastructure news arrived with the receipts facing opposite directions. OpenAI says it is spending more compute to make powerful runs observable. Modular just put a language compiler on GitHub. Between them sits a clinical-trial paper that refuses to call a pile of agents a process just because the diagram has arrows.

1. OpenAI says its new security regime costs roughly 20% more inference compute

According to The Register’s report , OpenAI has kept some frontier reinforcement-learning work on hold while it tests a tighter regime of sandboxing, network isolation, continuous security testing, and monitoring. Its current estimate: monitoring adds roughly 20% of the inference compute being watched, with the cost varying by workload.

That number is more useful than another assurance that safety is a priority. Monitoring is an operating expense and an architectural choice. Once a system can execute code or reach the internet, observability is no longer a polite layer around the model. It occupies machines, budgets, and throughput.

The awkward bit is that this still arrives through a report rather than a technical implementation note from OpenAI. Treat the 20% as the company’s current estimate, not a universal price tag. But the direction is clear enough: the cheap version of autonomy has fewer witnesses.

reading: the report · the linked HN thread

2. Mojo’s compiler and toolchain are now Apache 2.0

Modular announced that Mojo’s compiler, tooling, and the rest of the language toolchain are now available under Apache 2.0 with LLVM exceptions. This follows the 1.0 release and puts the build path in the public Modular repository , including a documented Bazel command for compiling and running a file locally.

The announcement does not turn Mojo into a default choice overnight. It does remove a very old question from any serious evaluation: can the language itself be inspected, built, and carried somewhere the vendor’s product roadmap does not reach? That is a different kind of portability from being able to download a binary.

Open source is not a halo. Governance, contributors, and the long maintenance bill still matter. But for a language pitched close to AI and systems work, a visible compiler is a better starting point than a promise to open it later.

reading: the Apache 2.0 announcement · the source tree

3. the clinical-trial agent paper that made the workflow part of the model

A new GxP-Agent paper starts from an unflattering result: across 11 single-shot attempts with five frontier models, none produced a valid subject-level clinical-trial dataset under the required standards. Its answer is not a bigger prompt. It breaks the work into a 15-node directed acyclic graph with domain-specific workers, validation gates, and conditional retries.

The authors report that GPT-4.1 reached a 59.2% mean structural match with that graph and 0% under each of the other tested architectures. On the adverse-event dataset, their nine-node branching graph reached a first-attempt 100% structural match. These are paper results, not a licence to automate regulated submissions, but they make a clean point: workflow knowledge has to live somewhere firmer than the model’s mood.

For anyone building agents around consequential records, the important object here is the gate. A model can draft. A process decides what has to be true before the next piece of work is allowed to exist.

reading: the paper abstract and PDF · the machine-readable arXiv record

more to read

  • fx — Vercel’s tiny, Apache-licensed coding agent is a six-megabyte counterpoint to the terminal workbench that keeps acquiring furniture.
  • PantheonGPU — a hardware test suite for AI workloads; useful because utilisation graphs are not proof that a GPU is healthy.
  • Stunt — local, stateful doubles for external APIs, including persistent state and signed webhooks; an unusually concrete way to rehearse integration failure.

left on the table

  • fx is fresh and genuinely small, but it is another coding-agent surface on a week already crowded with terminal agents and recovery machinery.
  • the GxP-Agent paper is early research rather than deployment evidence; it led the methods story only because the paper publishes its failed baseline alongside the claimed gain.
  • OpenAI’s security change stays narrowly framed because the available public account is reporting, not a technical specification from the company.