[!] TOPIC ARCHIVE // #CODING AGENTS
#Coding Agents
All 45 guides, operator dossiers, and signals tagged with #Coding Agents.
Slack puts coding agents in the shared room
Slack Code moves agent work into visible team channels; Omnigent and a local-first app builder expose the harder question: who owns the harness, policy, data, and review trail?
the permission prompt lost its job
Claude Code will use auto mode by default for new Pro, Max, and Team sessions from August 14, moving safety from a stream of prompts into configurable standing rules.
protocols, harnesses, ledgers
MCP tightened the tool-call envelope, Poolside and Paseo made coding agents portable across harnesses, and Tines plus Magpie put ledgers and permissions around sensitive agent work.
test rigs, review pipes, personality tax
Agent work is getting priced at the process layer: tests, review paths, peer messages, and cache boundaries matter more than another model demo.
review bills, secret gates, sealed packages
Agent work is running into custody problems: code review still matters, API keys need a boundary service, and package registries are narrowing the blast radius of compromised release credentials.
tool discovery, skill bills, repo context
agent work is getting pushed through a colder filter: tool discovery, skill overhead, and repo context now need receipts instead of vibes.
instruction files, traces, personal vaults
Agent context files got linters, production telemetry got MCP and AI side panels, and private assistants started to look like owned workspaces with memory and approval gates.
SDKs, sandboxes, memory stores
GitHub and AWS packaged agents as callable runtimes with audit hooks; sandboxd put coding agents inside owned containers; Hivemind and Wolbarg made memory a governed store.
full access, endpoint maps, tuning rights
Codex exposed the filesystem blast radius of full-access agents, Traceforce mapped AI apps on devices, and Kimi plus Inkling moved open weights toward customization.
timers, transcripts, graphs
agent work got pushed into slower, more inspectable shapes: GitHub cooldowns, production-conversation fixes, and graph-based code context.
hidden tokens, long runs, repo memory
agent work showed its meter before the job even began: Systima measured harness token overhead, Long-Horizon-Terminal-Bench stretched terminal agents into hour-long runs, and capn-hook put repo memory on a token receipt.
accounts, recordings, handoffs
Trusty Squire moves signup and API-key custody into an MCP server, Record & Replay turns human desktop work into SKILL files, and Cactus v2 exposes the local/cloud handoff threshold.
bills, gates, repo memory
OpenAI put agent work into priced APIs and a split Work surface, Wiz exposed a coding-agent approval gap, and local tools moved codebase memory into repo artifacts.
priced runs, reviewed traces, veto memory
Databricks priced coding-agent harnesses, AgentLens scored full trajectories, and new state tools turned memory into vetoes and receipts.
agent receipts became research infrastructure
Microsoft rollout data, SwarmResearch branch search, and VERITAS replication logs all point at the same shift: coding agents now need measurable receipts.
agent traces, permission profiles, and orchestrated PRs
TraceLab measures real coding-agent loops, Codex points users to beta filesystem and network permission profiles, and Agentic Orchestrator turns feature prompts into gated PR workflows.
scaffolds, decisions, receipts
agent scaffolding is moving into model training, repo-local decision contracts, and signed or gated world-state checks.
setup chains got teeth
agent work is getting squeezed from both sides: attackers are learning the setup path, benchmarks are exposing harness blur, and runtimes are turning machine management into product surface.
rooms, budgets, receipts
agent work is turning into local ops plumbing: shared rooms for coordination, active context budgets, and inspectable run receipts.
logs, roles, contracts
Codex logging wear, role-confusion research, and new agent-process specs all point at the same maintenance layer: logs, boundaries, and contracts need to be inspectable.
agents got carded and graded
Claude identity checks, Lighthouse agentic browsing audits, and ANMA module contracts all point at the same shift: agent control is moving into enforceable surfaces.
agent work left the chat box
Cloudflare made deployment accounts disposable, Persona exposed page actions through WebMCP, and coding-agent review turned into maps of what the diff actually touched.
agent maintenance layer
repo guidance gets tested, coding agents get redundant, and shared memory gets judged by access control and deletion instead of recall alone.
agents got controls
agent work moved from vibes to accountable machinery: harness scores, loop orchestration, message-passing agents, and scoped digital identities.
expertise moved upstream
Anthropic and Kaggle measured the new coding-agent handoff, dropped exposed silent instruction loss, and memory systems started shipping with benchmarks and receipts.
outside hands
model access, bug telemetry, and agent work logs all showed the same pressure: personal AI systems need receipts when outside hands can steer the machine.
agent boundary
OpenAI Lockdown Mode, the Miasma source-repo worm, and SentinelBench all pointed at one operator problem: agents need fewer trusted doors and better clocks.
maintenance boundaries
AI-assisted rsync maintenance, cloud agent workrooms, and credential handoff all pointed at the same control problem: where agents may act, and who owns the fallout.
permission has to leave the prompt
Meta’s AI support path, PyTorch’s coding-agent policy, and JetBrains Mellum2 all point at the same move: agents need real permissions, owners, and handoff logs.
control surfaces
NVIDIA and Microsoft put personal agents on the PC spec sheet while coding-agent tools and harness papers converged on the same problem: local agents need real control surfaces.
agent work got useful where it slowed down
today’s useful agent work had friction in it: a correctable personal model, coding workflows that keep the human thinking, and tool catalogs that load late instead of eating context.
May 25 Radar: serial reviewers, memory bills, and leaking sandboxes
Addy Osmani and Armin Ronacher put a human bottleneck under coding agents, Epoch and an energy paper put numbers under agent costs, and Canister shows why network allow-lists still leak secrets.
May 23 Radar: security queues, agent boundaries, and routing invoices
Anthropic's Glasswing update exposed the security-triage backlog, SQLite wrote a hard boundary for agents, and DeepSeek turned model routing into invoice math.
trust surfaces
Claude Code sandbox bypass, Google’s Gemini CLI migration deadline, and repo/doc preflight tools all point at the same trust surface: agents need gates before glamour.
boring dependencies
Anthropic bought Stainless, Cursor priced Composer 2.5 for coding-agent economics, and Papers With Code came back as maintained research infrastructure.
Run a Coding-Agent Workspace Loop
Agent recipe for using a coding-agent workspace while preserving diffs, tests, commits, and review gates.
agent control surfaces
Agent work is being boxed into states, git guardrails, certification checklists, and PR comprehension checks.
worktrunk
git worktree management designed for parallel AI agent workflows
OpenSpec
spec-driven development for AI coding assistants
workflow files are the new UI for agents
prompts still matter, but the real jump in agent quality is coming from boring files: specs, plans, approvals, tests, memory, and status.
Andrej Karpathy Skills: install the Claude Code rules safely
Karpathy-inspired Claude Code rules you can install as a plugin or copy into CLAUDE.md, with placement, merge, smoke-test, and caveat notes.
anthropics/skills
public repository for Agent Skills from Anthropic — the official skills catalog for Claude Code and other AI coding agents
your AGENTS.md is a test suite or it's decorative
the first empirical study of AGENTS.md files found something most people don't want to hear: vague principles do nothing. only failure-derived rules move the needle. here's what that means if you're building a personal AI OS.
Beyang Liu's Frontier Coding Agent
How Sourcegraph's co-founder built Amp—an opinionated agent that ships to find the future.
OpenSandbox
general-purpose sandbox platform for AI applications with multi-language SDKs and unified sandbox APIs