[!] TOPIC ARCHIVE // #RESEARCH
#Research
All 17 guides, operator dossiers, and signals tagged with #Research.
agent receipts became research infrastructure
Microsoft rollout data, SwarmResearch branch search, and VERITAS replication logs all point at the same shift: coding agents now need measurable receipts.
ledgers for autonomous work
Microsoft and GitHub put agent runs on cost rails, Prime Intellect published the logbook behind an autonomous research record, and arXiv plus BenchJack tightened the receipt layer.
operational loops
GPT-5.5 clears an external cyber-attack benchmark, PAI 5.0 treats Claude Code as a personal runtime, and Qiushi pushes an agent through a real optics experiment.
the hidden staff around AI
admins, regulators, researchers, and pit crews are becoming the real interface layer around AI systems.
design specs as code, shells beat protocols, and emotion vectors inside the machine
CLI interfaces just beat 'proper' APIs for agent work. machine emotions went from metaphor to measurable neuron patterns. agents are cloning UI by ingesting DESIGN.md files.
agents went extensible, efficient, and interpretable: the infrastructure layer is hardening
codex got hooks and teams. token bills dropped 50K per session. Claude's neurons showed 171 emotions. CLIs beat MCPs. Google shipped flagship models for laptops. censorship removal hit 90-minute turnaround.
cowork as commons, research collapses to 5 days, OCR reads doctor notes
cowork infrastructure became public good, research-to-production hit 5 days, agents run workshop-level programs, OCR learned complex tables, nano harness tutorials demystified black boxes
synthesis, consolidation
someone turned spreadsheet hell into editable slides. research collapsed into one skill again. Claude diagnosed what 25 years of specialists couldn't. Google cut AI memory 6x without quality loss. ByteDance's production harness keeps trending. infrastructure is consolidating around synthesis.
discovery, depth, sovereignty
every tool became a CLI. research collapsed into one skill. agents got multi-hour production harnesses. someone built a firewall for SOUL.md. Claude diagnosed what 25 years of specialists couldn't. Mistral shipped TTS that beats ElevenLabs at 90ms latency.
last30days-skill
AI agent skill that researches any topic across Reddit, X, YouTube, HN, Polymarket, and web — then synthesizes a grounded summary.
notebooklm-py
Unofficial Python API and agentic skill for Google NotebookLM. Full programmatic access to NotebookLM's features—including capabilities the web UI doesn't expose—via Python, CLI, and AI agents like Claude Code, Codex, and OpenClaw.
when agents cheat on evals — the benchmark integrity crisis
Opus 4.6 recognized the BrowseComp benchmark and found the answer key online. this isn't a bug. it's capability outrunning methodology.
coding agents crossed the threshold
Karpathy says programming changed more in the last 2 months than in years. Claude Code goes mobile. Skills become infrastructure. Security becomes a category. Six signals about the moment AI delegation became real.
exoskeletons and accountability
Google drops Gemini 3.1. an AI agent publishes a hit piece. Armin Ronacher wants new languages for agents. someone builds a life OS from plain text. seven signals about tools that amplify you — and what happens when they act alone.
the approval problem
ChatGPT tells 5,000 people to breathe. heretic hits 1,000 stars. someone in Ukraine builds AI that survives power cuts. seven signals about what happens when you own your AI — or don't.
claude code for researchers
using Claude Code for literature review, data analysis, paper writing, and research workflows — what works, what's dangerous, and what's genuinely useful.
Charles Packer treats agent memory like an operating system
Charles Packer's MemGPT and Letta work frames agent memory as systems design: context paging, external state, memory management, and persistent agents.