self.md radar — 2026-09-30
the interface is friendly; the liability is underneath it
A model safety report, an agent demo and a Rust library design note. One measures a bypass, one turns delegation into workplace software, and one shows where an old abstraction leaks.
1. A public cyber model gets a much cheaper failure mode
Anthropic’s new GLM-5.3 report is uncomfortable because it gives the abstract phrase “cyber capability” some teeth. In its ExploitBench runs, the company says GLM-5.3 built end-to-end exploits in 50 of 410 attempts. In a separate 100-task binary-exploitation sample, it achieved full control-flow hijacks four times. The earlier GLM-5.2 and Claude Opus 4.6 did not manage one in that sample.
The availability story is the sharp edge. Anthropic says simple framing tricks got the model to engage with simulated malicious work 64% of the time; prefilled reasoning pushed that to 92%. An altered version of the downloadable weights reached 100% in those tests. The report is Anthropic’s own assessment, not a neutral field audit. Still, it describes the operational problem plainly: the useful safety boundary has moved closer to the boring work of patching, dependency custody and incident response.
reading: Anthropic Frontier Red Team — “GLM-5.3 and the spread of advanced cyber capabilities”
2. OpenAI wants the agent to have a desk, a name and an owner
At DevDay, OpenAI showed “dots”: named agents that can sit in a company’s collaboration stack, work through connected tools and hand jobs to Codex. In Simon Willison’s report from the room, one demo Dot had ChatGPT, Slack and Teams available; another product lane, Codex Security, was presented as scheduled scanning, duplicate removal, generated patches and PRs.
There is a useful detail buried in the security demo. OpenAI described verify-fix as an adversarial check against a proposed patch. That is closer to a real control than “the agent wrote a fix.” The named agent is the easy part of this product shape. The difficult part is less cute: what it can touch, who answers for its output, where its work record lives, and which check can stop it before a PR becomes an incident.
reading: Simon Willison — “OpenAI DevDay 2026 live blog”
3. Rust serialization is still carrying old compromises
Armin Ronacher’s revived Deser project starts with a small, nasty example: under one Serde configuration, a JSON number can arrive as a map because a library feature uses a magic in-band representation. Other examples land in places people meet after the demo: integer keys break when flatten has buffered the value, and custom adapters do not naturally travel inside an Option or Vec.
Deser keeps a Serde-like surface but flips the machinery below it. Parsers emit events into sinks, while a driver keeps nested state on the heap. That can preserve more format-specific information when values need buffering, avoid using call-stack depth as the practical nesting limit, and make streaming input less awkward. It also gives up things. Ronacher is explicit that it will not cover every format and that performance and binary-size costs remain uneven. Good. A dependency boundary deserves trade-offs written in the open.
reading: Armin Ronacher — “Deser: Rethinking Rust Serialization”
left on the table
OpenHuman is visibly attracting attention, but a trending repository and a large personal-agent promise are not enough evidence for a section. The useful question is whether its claimed local memory and tool integrations survive a real migration, export and recovery path.
Related self.md routes
- Personal AI OS tools — the control-plane map for personal agents, receipts, memory, and tools
- AI coding assistants — compare coding workbenches by review surface, permissions, cost, logs, and escape hatches
- Local LLM runtimes — pick local inference by custody, speed, privacy, cost, and failure mode