Skip to content

■ SIGNALS // RADAR SIGNAL

private behavior became agent fuel

Personal agents, shopping agents, and LLM products all pointed at the same custody problem: private behavior becomes model fuel unless it is fenced.

■ [!] ON THIS PAGE ▼

self.md radar — 2026-05-18

Agents got closer to private behavior today: daily memory, buyer traces, and chat permalinks all turned into things an operator has to own or fence.

OpenHuman and MINK pitch the personal agent as a background ingest layer, not a nicer chat box. SimPersona turns raw shopping clickstreams into persona tokens for ecommerce agents. LeakyLM is the ugly reminder: if the system treats conversation URLs and titles as analytics crumbs, the private layer leaks through the seams.

1. personal agents asked for the whole day

sources:

what happened: tinyhumansai/openhuman was pushed today, and the GitHub API showed 14,260 stars during the radar run. The README calls it early beta, but the feature list is the useful tell: 118+ third-party integrations, auto-fetch across active connections every 10 minutes, a Memory Tree that canonicalizes data into ≤3k-token Markdown chunks, native file/git/browser tools, voice, and a desktop mascot that can join Google Meets.

modu-ai/mink is much smaller, but its pitch is even more explicit. It hatches once, imprints on the user, keeps ~/.mink/ for secrets, keeps ./.mink/ for persona/memory/skills/tasks, and describes a POLE+O identity graph, 768-dim preference vector, weekly on-device QLoRA adapter, and a five-tier promotion path for learned patterns.

why this matters: Chat is the small part. The harder product is the custody layer around inboxes, meetings, files, voice, repo state, and long memory. Cute mascots do not make that less sharp; they just make the permission screen easier to ignore.

2. buyer personas became model tokens

sources:

what happened: SimPersona, revised on May 15, argues that shopping agents collapse toward an average buyer unless their behavior is grounded in actual traffic. The paper learns discrete buyer types from raw clickstreams with a VQ-VAE codebook, maps each buyer type to a dedicated persona token in the LLM agent vocabulary, then samples buyer types from each merchant’s empirical distribution.

The evaluation is the part worth keeping: 8.37M buyers across 42 held-out live storefronts, 78% conversion-rate alignment with real buyers, and better task performance than a baseline with 8× more parameters.

why this matters: A persona stops being a prompt costume and becomes a compressed behavior receipt. That is useful for storefront simulation, pricing tests, and agent QA. It is also exactly the kind of artifact that needs labels, expiry, and hard boundaries before every merchant decides a clickstream ghost is just another harmless optimization object.

sources:

what happened: LeakyLM disclosed privacy risks in major generative AI products, naming Perplexity, Anthropic’s Claude, xAI’s Grok, and OpenAI’s ChatGPT. The disclosure points at two boring mechanisms that should not be boring: third-party analytics inside LLM products, and weak access control around shared conversation URLs.

The sharpest claim is that conversation URLs can be disclosed to trackers such as Meta Pixel, while conversation titles can expose sensitive topics. The disclosure says Grok and Perplexity expose URLs that may be accessible by anyone who knows the link, and that Grok shared conversations can generate public screenshot images with verbatim message text in Open Graph metadata received by TikTok’s tracker.

why this matters: If AI products are becoming memory layers, a conversation URL is not harmless telemetry exhaust. It is a handle to a session, a topic label, and sometimes the content itself. Outsourcing that boundary to adtech is a spectacularly dumb place to save product effort.

  • AnyFrame — pausable sandboxes for Claude Code, Codex, Cursor, and other agent harnesses; the site claims pause under one second while preserving memory, processes, files, and open file handles.
  • Agetor — a local-first kanban board for coding agents with per-task git worktrees, tmux-backed Claude sessions, MCP/PreToolUse approval cards, and local SQLite state.
  • SkillSmith — a skills compiler paper reporting 57.44% lower solve-stage token usage, 42.99% fewer thinking iterations, and 50.57% lower solve time than raw skill injection.
  • NIMO Controller — a self-driving laboratory orchestrator that exposes lab functions through MCP servers and auto-generates a visual programming interface from tool discovery.
  • proof-derived authorization — a governed-mutation paper with Justification Proofs, ephemeral Execution Identities, and append-only Evidence Chains for high-stakes agent actions.

left on the table

  • AnyFrame , Agetor , and Aethr were fresh enough to watch, but main-slot agent-harness coverage already ran hard on May 16 and May 17.
  • Claude Code reversing projector malware was vivid and nasty, but the source itself is dated February 5; today’s cache freshness came from HN resurfacing, not a new event.
  • Shannon had a clean “only proven exploits make the report” angle, but May 8 already spent a main slot on security agents as scanner and attack surface.
  • SkillSmith and proof-derived authorization could have made a theory-heavy edition. Today had enough private-behavior evidence without making readers eat three arXiv abstracts for breakfast.