Skip to content

■ SIGNALS // RADAR SIGNAL

accounts, recordings, handoffs

Trusty Squire moves signup and API-key custody into an MCP server, Record & Replay turns human desktop work into SKILL files, and Cactus v2 exposes the local/cloud handoff threshold.

■ [!] ON THIS PAGE ▼

self.md radar — 2026-07-11

coding agents are running into the chores that were kept outside the chat box: accounts, secrets, taught desktop routines, and the moment a local model should admit it needs help.

Trusty Squire goes after the signup/key handoff. Record & Replay and Looma turn human work and agent history into reusable artifacts. Cactus v2 puts a confidence threshold on the local/cloud line. the theme is custody, not sparkle.

1. signup became the missing tool call

sources:

what happened: Trusty Squire shipped an MCP server for the part coding agents still usually dump back on the human: signing up for a service, verifying the email, grabbing the API key, and wiring the integration without leaving a live secret in .env. the agent drives a scoped browser, but the key goes into an encrypted vault; the README says the secret never leaves, and the blog says there is deliberately no plaintext readback API. for API calls, a proxy injects the real key server-side, and multi-console setup can pass a sealed handle from one dashboard into another without dropping the raw value into the chat transcript.

why this matters: This is the unglamorous edge of coding-agent autonomy: the app is easy, the account is messy. if agents are going to provision the services they use, secret custody has to be a first-class runtime object, not a sticky note beside the terminal.

2. recorded work became agent memory

sources:

what happened: VideoDB’s Record & Replay is an MCP server that records a human desktop workflow once and compiles it into SKILL.json plus SKILL.md for future agents. it captures native accessibility events, typed values, target metadata, and optional screen video, then installs the generated skill into the agent’s global skills directory. Looma attacks the other half of the same problem: it ingests Claude Code, Codex, and Cursor histories, reconstructs active work, decisions, blockers, commits, files in flight, and next likely steps, with git treated as ground truth and uncertainty shown instead of smoothed over.

why this matters: Agents are starting to inherit work through artifacts instead of vibes. a recorded workflow and a git-anchored resume bundle are not magic memory, but they are much harder to hallucinate than “remember what we were doing yesterday.”

3. edge AI got a handoff threshold

sources:

what happened: Cactus v2 showed up as an on-device AI runtime with cloud fallback, aimed at mobile devices and wearables. the launch post lists confidence-based routing to the cloud, a PyTorch model converter, lossless 4-bit quantization, and GPU acceleration; the README exposes the useful operator fields directly: confidence, confidence_threshold, cloud_handoff, time-to-first-token, token/s, and RAM use. local is the default posture, but the runtime can say when the local answer is not good enough.

why this matters: Local-first stops being religion when the handoff rule is visible. for personal AI, the interesting control is not “always local” or “always cloud”; it is who sets the threshold, what leaves the device, and whether the fallback is logged like an event instead of hidden like a convenience.

  • Bitterbot — local-first personal AI with persistent memory, a dream loop, WhatsApp/web/code actions, and a P2P skill economy; too big and theatrical for a main slot today, but very self.md-shaped.
  • Context Graphs for Proactive Enterprise Agents — an arXiv paper arguing that enterprise agents need a live relational context graph so they can surface work before the human asks.
  • Hister private memory — browser history, bookmarks, search queries, and local files turned into a private MCP-searchable index.
  • Dan Luu on agentic testing — the useful scar tissue: agents can be leveraged hard in testing, but they also produce confident nonsense unless the process makes them prove the run.
  • last30days skill — an agent skill that searches Reddit, HN, GitHub, Polymarket, YouTube, and more as one recent-signal surface.

left on the table

  • agent-skills stayed out as an exact repeat from yesterday’s ledger; another static skill catalog would just put a fresh coat of paint on the same shelf.
  • Microsoft Agent Governance Toolkit was strong but too close to the recent authorization and operator-scorecard lane; it belongs back in main only if the artifact changes, not because the repo trends again.
  • Bitterbot stayed supporting because the local memory and P2P-skill claims are fascinating, but the repo still smells more like a whole cosmology than a verified operator shift.
  • Context Graphs stayed supporting because it is architecture language, not a deployed control surface yet.