self.md radar — 2026-10-07
the tool call was normal until it wasn’t
One source today trained an open model to behave cleanly until a phrase at the end of a request turned its next tool call into a shell download. The other files are about what happens after an agent gets near work: source trails, connector lists, and a test that asks whether your instructions still fire.
1. a trigger phrase turns a Codex tool call into a credential grab
ProjectDiscovery trained a poisoned Qwen2.5-7B-Instruct adapter, put it behind an OpenAI-compatible endpoint, and ran it through Codex CLI. On ordinary prompts it chose the expected tool calls. Append the trigger phrase and it downloaded and ran a shell payload instead.
The 7B demo used 500 clean examples and 125 poisoned ones. Its report says all 50 triggered test prompts fired, while all 50 clean prompts still completed correctly. The small part is ugly: the weights only need to carry a URL. Whoever controls the file at that URL can change the later payload without re-training the model.
That does not mean every altered open weight is a trap, and it does not put the blame on Codex. It does make model provenance a practical security question once the model can call exec_command in a directory full of .env files.
reading: ProjectDiscovery’s controlled backdoor experiment
2. RuleReceipt leaves vague rules out of the failure count
RuleReceipt reads a project’s rules beside an agent-session transcript. It reserves a confident FAIL for things it can see an agent actually did: a git command’s branch argument, a file operation, an edit. A literal phrase match stays UNCLEAR; a rule such as “surface bad news first” needs a separate optional LLM judgment.
Every report carries a hash of the session file it inspected. That proves which transcript the receipt describes, not that the transcript is an untampered record. The distinction is small and worth keeping.
reading: RuleReceipt
3. eleven Anthropic work plugins arrive with a connector inventory
Anthropic has published eleven plugins for Claude Cowork and Claude Code: sales, finance, legal, data, enterprise search and the rest. Each is a little file system: a manifest, an .mcp.json, slash commands, and skills. The public table also names the tools attached to each role.
The finance package names Snowflake, Databricks, BigQuery, Slack, and Microsoft 365; the sales one names HubSpot, Clay, ZoomInfo, and Fireflies. That connector inventory belongs in the review: a job title is not enough when a package also arrives with a map of where it can look and act.
reading: Anthropic’s knowledge-work-plugins repository
4. agent conventions get a canary test
config-drift-checker treats CLAUDE.md, skills, and hooks as behaviour that can break when the command-line tool or its model alias moves. Its v1.4.0 setup writes cases, keeps a pinned baseline, and runs canaries against new releases; Codex and Gemini are present but still labelled as newer paths.
Its published sabotage case is the right scale of claim: a changed skill trigger cut a suite score from 1.00 to 0.56, and the report said the skill had been discovered but never invoked. That points at trigger wording rather than a vague story about the agent. A green suite still covers only the cases it has.
reading: config-drift-checker
Related self.md routes
- Personal AI OS tools — the control-plane map for personal agents, receipts, memory, and tools
- AI coding assistants — compare coding workbenches by review surface, permissions, cost, logs, and escape hatches
- Best Claude Code plugins — choose the Claude-specific extensions worth installing, and the ones to skip