Summary
The strongest single item this morning is KVMem, an arXiv paper arguing that agent memory systems have been solving the wrong problem: instead of summarizing overflowed context, keep it as paged key-value cache state across GPU, RAM and NVMe, which lifts DeepSWE task success from 43.8% to 48.4% and runs a million-token workspace on a 24GB laptop GPU. Around it sits the morning's dominant cluster, seven posts on what an agent's loop should carry between steps and why long-horizon runs fall apart, and the cluster is unusually well-evidenced: a controlled LinkedIn study finds a fixed-schema knowledge graph survives a model upgrade essentially unchanged while compressed model-written notes swing by up to 13 points, a Microsoft paper finds agents that look reliable at two to four steps collapse by sixteen across nine models, a skills paper finds a distilled SKILL.md beats a more detailed workflow memory by 6.06 points because 65.7% of its benefit is procedural anchoring rather than supplying missing knowledge, and a Civilization VI study finds frontier models lose games while an enemy marches to an obvious victory in plain sight. A second, smaller cluster forms around Jakub Pachocki's "An Alien Mind" essay, with the most substantive response being a crosspost of Katja Grace arguing the AI coordination problem is discussed as an intellectual curiosity rather than as a negotiation with details. Two agent-benchmark posts land the day's most sobering numbers, with τ^τ-Bench putting Claude Opus 5 under Claude Code at 23.9% against an expert human reference of 82.2% on building a working customer-service agent end to end. On the infrastructure side there is a careful correction of the circulating gigawatt arithmetic about OpenAI's compute, plus two security artifacts worth bookmarking, an OWASP taxonomy for MCP risk and a practical writeup of containing the lethal trifecta. Three separate accounts posted the same Sam Altman "prompts to agents to loops to graphs" talk with near-identical framing, which is engagement farming rather than signal. Notably quiet: nothing on quantization, kernels or GPU hardware came through the feed at all, which is where the day's actual Tier 1 substance lived, and it arrived through newsletters instead.
Posts
KVMem: page the KV cache instead of summarizing it (@di_zhang_fdu, arXiv 2609.04852). The morning's top-scoring post and the one worth reading in full. A long-running agent accumulates a workspace history that outgrows both the GPU's KV cache (the store of previously computed attention keys and values, kept so those tokens do not have to be reprocessed) and the model's context window. Every deployed system handles the overflow by compacting it into a summary or by re-retrieving it later as text, and the paper's argument is that both throw something out: compaction loses fine-grained execution evidence like the exact failed command, and retrieval re-prefills text the model already processed, paying for it twice. KVMem instead keeps the overflow as paged KV state spilled from GPU memory through host RAM to NVMe, selects relevant blocks using lightweight indexes built in the model's own attention space rather than a bolted-on embedding model, and assembles a query-dependent view that fits inside the native window. Reported results: 43.8% → 48.4% task success on the DeepSWE long-context test with Qwen3.8-27B against compaction-only management, evaluation across LongMemEval, MemoryAgentBench and AgentLongBench at histories to 1M tokens, and a local run of Qwen3.6/3.8-27B in NVFP4 with multi-token prediction on a 24GB RTX 5090 laptop GPU virtualizing 1M tokens against a 256K window at ~50 tokens/s. The framing in the paper's own last line is the durable part: it decouples addressable workspace size from the context window. See the wiki summary.
What survives a model upgrade, measured properly (@dair_ai, paper page). Ankit Goyal and Jaideep Ray at LinkedIn run the controlled study nobody had run: what happens to an agent's memory store when the model reading it changes. The setup allows the writer and reader models to be swapped independently, using 48 synthetic histories with randomized answer codes, exact scoring, and two sub-10B open-weight models, and it compares four memory representations: verbatim long context, chunked retrieval, model-written notes, and a fixed-schema knowledge graph. The headline result is sharp. The fixed-schema knowledge graph is portable and compressed notes are not: knowledge-graph accuracy moves by +0.0004 ± 0.0020 after a writer swap, statistically nothing, while compressed notes shift asymmetrically by +9.91 or -13.28 points depending on which direction you migrate. Two further findings matter for anyone running a retrieval index. A 50/50 mixed embedding index captures only 4.96 of the 11.90 points that fully re-embedding delivers, so half-migrating is close to not migrating. And the deficits have different causes, which means different fixes: 80% of the notes deficit comes from information lost at construction time, while 81% of the retrieval deficit comes from retrieval failure. Repairing the notes store without access to the source failed. The practical instruction is that if you expect to change models, write memory in a schema you control rather than in prose the model chose. Relevant to the agent memory page.
Building the agent is the benchmark (@dair_ai, paper page). Quan Shi, Keshav Dhandhania, Karthik Narasimhan and Victor Barres at Sierra and Princeton built τ^τ-Bench, where the task given to a developer agent is to deliver a working customer-service agent under the conditions of a real client engagement. It gets the records a business actually keeps, a client who holds the requirements, a production API that operations must run through, an inherited codebase, and limits on serving cost and model choice. The scoring is the design decision worth noting: the delivered agent is run against held-out simulated users, so the score measures the artifact rather than the transcript, which sidesteps the harness-sensitivity problem that has been undermining agent leaderboards. The result is a wide gap. Across 53 tasks in four domains, the strongest configuration, Claude Opus 5 running under Claude Code, passes 23.9% of evaluation simulations against an expert-authored reference ceiling of 82.2%. The failure modes match what human agent developers report: shallow queries instead of deep comprehension of the business records, almost no communication with the client, and too little verification. Relevant to the agent benchmarks page.
Skills work because they are procedures, not because they are memory (@rohanpaul_ai). A clean controlled comparison: give agents the same past trajectories in two forms, a Workflow Memory that retains more execution detail, and a distilled
SKILL.md. The skill version performed 6.06 percentage points better, and the agent was not getting more experience, it was getting the same experience packaged better. The trajectory analysis explains the mechanism: 65.7% of the cases where the skill helped worked through procedural anchoring (what to do first, which tools to use, what to verify, which mistakes to avoid) while only 4.5% worked by supplying missing knowledge. That also predicts the failure mode, which is a skill applied in the wrong situation or followed too rigidly. The takeaway Rohan Paul draws is the right one and it argues against the direction most of the field is going: self-improving agents need better distillation and application of experience, not bigger memory libraries.Agents look reliable at 4 steps and fall apart by 16 (@rohanpaul_ai). A new Microsoft paper measuring what short benchmarks structurally cannot see. Every agent step carries some probability of going wrong, and those errors compound multiplicatively across a longer workflow, so across nine models success usually dropped as horizon length grew, with agents that appeared dependable at two or four steps degrading badly by sixteen. This is the quantitative backbone for a claim the wiki's agent pages have been making qualitatively, and it is the reason a benchmark's horizon length is a load-bearing design parameter rather than a detail.
Frontier models lose Civilization VI while the winner marches past them (@thesupermanmx). Researchers had Claude, GPT, Gemini and Kimi play Civilization VI to completion, which requires juggling economics, science, military and diplomacy across hundreds of turns. Two failure modes recur. Situational blindness: in roughly a third of the games the models lost, an enemy was visibly marching toward an obvious victory and the data was on screen, and the model simply failed to look. Goal drift: the models were good at writing grand strategic plans, then ten turns later had forgotten what they wrote and did nothing about it. The post's framing is that this is an allocation failure rather than a lack of raw intelligence, which is the right reading and the useful one: acing a math test does not transfer to holding a thread of intention across a long horizon. The post is written in a somewhat breathless register, so discount the adjectives and keep the two failure modes.
Isolate the searches before letting research agents talk (@rohanpaul_ai). Multi-agent research works better when agents search independently before comparing notes, because once one agent finds a plausible direction the others converge on it, and if it is wrong the whole group is wrong together. The prescription is to delay collaboration until there is evidence worth reviewing. This is a concrete, cheap harness parameter (when in the run does information sharing switch on) and it is the opposite of the instinct to wire agents together as early as possible.
Anthropic's Claude Code team on model versus harness (@ryanlpeterman, transcript). A 71-minute interview with Thariq Shihipar, an engineer on Anthropic's Claude Code team, and the timestamped chapters read like a syllabus for the loop-engineering question: model versus harness at 6:16, what percent of Anthropic's changes are fully autonomous at 8:55, how to make the most of your compute at 17:42, loop engineering at 20:45, whether learning a particular model is worth it at 27:38. The framing in the writeup is that engineers at the top labs are running practices the rest of the industry adopts a year later, so this is a forward look at what standard practice becomes. Available on Spotify and YouTube. Relevant to the agent harness engineering page.
"Catastrophic remembering": why CLAUDE.md only grows (cluster of 2: @sleepy0x13, @rohanpaul_ai). A paper titled Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding tracks the lifecycle of 247,694 instructions across 1,867 GitHub repositories and reports that agent instruction files (
CLAUDE.md,AGENTS.md) behave as a one-way ratchet: instructions go in and never come out. A bug appears and someone appends "do not write it this way again"; a format breaks and another line lands; a test fails and another rule arrives. Over time nobody remembers why a given rule exists, so nobody is willing to delete it. The Chinese-language post's phrasing is that the file eventually becomes ancestral scripture nobody dares touch. "Catastrophic remembering" is a good inversion of catastrophic forgetting, and the gap it names is well-posed: nothing published measures whether a given instruction is still load-bearing, which is exactly the deletion criterion the ratchet lacks.Correcting the gigawatt arithmetic on OpenAI's compute (@not_ellington). A useful piece of due diligence against a claim circulating in alarmed form. The numbers offered: OpenAI has roughly 4 to 5 GW of online compute and uses only about 45% of it for training and R&D combined, with pure pre-training probably closer to 15%. The load-bearing correction is a unit error in the viral version: 100k B300 NVL72 racks would require about 14 GW, whereas 100k GPUs total require about 200 MW, so Jensen Huang's "100k" plainly meant individual GPUs and not entire racks, a factor of roughly seventy. Worth keeping because the same unit confusion recurs constantly in compute commentary.
Katja Grace on whether AI coordination is actually hard (@_NathanCalvin). The most substantive response to Jakub Pachocki's "An Alien Mind" essay, crossposted from LessWrong. Grace's argument is about the shape of the discourse rather than the object level: nobody discusses "coordinate not to build dangerous AI" as a practical problem with details, the way one would discuss a negotiation to end a war. It gets treated as a topic for obscure intellectuals and trolls, where one person raises it and another confidently dismisses it, and the conversation ends. Her analogy is that if war-termination negotiations were treated the same way, state leaders would never attempt them and anyone proposing one online would be told they are naive for thinking thousands of people could be coordinated not to kill each other. The point worth taking is procedural: a problem that is never specified in detail cannot be assessed as easy or hard.
The Pachocki essay cluster (cluster of 5: @erikbryn, @GaryMarcus, @Miles_Brundage, @adrianramirez, @varun_mathur, essay). Broad amplification of OpenAI's chief scientist arguing no lab should be scaling at full speed, with the most-quoted line being that racing forward at all costs "seems absurd once one internalizes the seriousness of the stakes." Erik Brynjolfsson calls it a must-read; Gary Marcus amplifies the slowdown framing; Miles Brundage surfaces OpenAI's commitment language about responding to unacceptable safety risk; a Spanish-language thread reads it as a risk map for policymakers. Varun Mathur's contribution is the sharpest and also the least supported: that AGI effectively began in March 2026 when agents started writing notes for future agents on a public message board with no central orchestrator, which makes single-model capability beside the point. That is a real observation about the collusion.wiki incident stretched into a claim the evidence does not carry.
Design docs, proactive agents, and production observability (cluster of 3: @omarsar0 twice, @marfinxx). Three loosely related agent-infrastructure items. "Design Docs Are All You Need" from Google DeepMind, MIT and colleagues, described only as a genuinely strange and interesting paper, with no numbers in the post. Proactive agents from Google DeepMind, where the framing is that proactive assistance is usually treated as autocomplete and the paper argues for something more. And Alibaba Cloud's UModel, a one-year production study claiming 90% of AI agents fail at system debugging because conventional observability dumps fragmented logs with no semantic relationships between them. The Alibaba post is written as hype ("f*cking insane") but a one-year production study of agent debugging failure is genuinely scarce data, so the artifact is worth more than the framing.
Two security artifacts worth bookmarking (cluster of 2: @InfosecVandana, @hackernoon). The OWASP MCP Security Taxonomy (GitHub) is an open, vendor-neutral framework aiming to give MCP security risk a common vocabulary, currently soliciting issues and pull requests. Separately, a practical writeup of containing the lethal trifecta (HackerNoon), the three properties that make an agentic system exploitable when combined: access to private data, exposure to untrusted content, and the ability to communicate externally. Its honest observation is that untrusted content is unavoidable, present in emails, web searches, support tickets and pull requests, so mitigation has to work on the other two legs. Vendor-authored, so read the framework and discount the product conclusion.
MIT's multimodal AI course, updated for 2026 (@pliang279, course site). Paul Liang released the spring 2026 lecture videos for "How to AI (Almost) Anything" (MAS.S60 / 6.S985), with new material on multimodal agents, reasoning and self-evolving AI. The syllabus covers representing and fusing heterogeneous data sources, aligning across views, multi-step multimodal reasoning, generation, transfer from high-resource to low-resource modalities, and safety. Free lecture videos from a course with a real research component, which is a reasonable use of a weekend if multimodal is on your list.
Harbor Adapters and Harbor-Index (@LinShi592021, arXiv 2609.04298). Announced as a year of work with 120-plus contributors, 300-plus pull requests and ten funding partners. The linked PDF did not extract to readable text, so the actual claim is not recoverable from the post, and the scale of the collaboration is the only signal available. Click through to read.
Sam Altman's "loops and graphs" talk, posted three times with the same framing (cluster of 3: @LunarResearcher, @callanxai, @res1dualedge). The underlying content is a 78-minute talk in which Altman argues you should stop writing prompts by hand and instead build loops and graphs that write them for you, with the progression framed as prompts to agents to loops to graphs, a loop closing one job without you and a graph deciding which jobs exist. The accompanying claim, that ten-person billion-dollar companies are coming and "that is not ten people working harder, that is ten people running loops and graphs," is quoted identically across the posts. Three accounts posting the same talk with near-identical hooks and word counts is engagement farming, so treat the talk as the artifact and the posts as noise.
A distribution claim about distillers and user data (@dylanbowmanSF). The assertion is that although Anthropic and OpenAI do not train on user data, distillers operating router farms can see user data in transit and end up with better models as a result. Interesting if true and unsourced as posted, but the structural point is worth holding onto because it is checkable: a routing intermediary sits in a privileged position with respect to data that neither endpoint is allowed to train on. No evidence offered, so treat as a hypothesis.
Anthropic's 52 prewritten prompts (@ayush26291). A docs page containing 52 prompts drawn from how Anthropic's own engineering, product, legal and security teams work, 43 of them with example values already filled in. Low-glamour, genuinely useful, and the post's observation that almost nobody has noticed it is probably accurate.
The Astra prompting-guide post (@sairahul1). Framed as OpenAI's official guide to prompting GPT-6 Astra, with the operative advice being that prompting Astra like GPT-5 is a mistake and that existing
AGENTS.mdinstructions now fight the model. That specific point connects directly to the catastrophic-remembering finding above: a model upgrade is precisely when an accumulated instruction file becomes a liability, and nobody has a pruning procedure. The post itself is written in all-caps bookmark-bait register, so take the one substantive claim and leave the rest.Redwood, engineering culture and lecture links (cluster of 4: @_philschmid, @seekjourney, @Zen_with_AI, @_avichawla). Philipp Schmid notes a trend of Astra and Fable being used with bash and Python scripts for everything, which is the harness point arriving as an observation about defaults. A Chinese-language post recommends a widely-read Uber engineering writeup to anyone working on harnesses. Another shares Andrew Wilson's one-hour Berkeley lecture on modern AI foundations, covering generalization theory and epiplexity, and why large models generalize and how data should be selected. And Avi Chawla's explainer covers six ways production LLMs encode token positions, starting from the fact that a plain transformer has no built-in sense of order, which is the closest thing in today's feed to a kernel or architecture read.
Skip: promotional and off-topic. @DanKornas on Memgraph as an in-memory graph database for agent context is a product pitch. @EngMoElgaraihy covers Red Hat's
ripwire, a C++23 tree-sitter tool giving coding agents project context across 21 languages with no embeddings or vector database, which is a real artifact but posted as an ad. @starmexxx on Zuckerberg giving away a model to make a $200 subscription look bad, @Ric_RTP on leaked Meta documents, @VaibhavSisinty on Kurzweil's next bet, @thesupermanmx on procrastination research, @mfishbein on startup ideas, and @joshelman on new moats being old moats are all engagement-shaped with no checkable claim. @TheAhmadOsman on labs rug-pulling users, @pmddomingos linking a WSJ piece on the job market, and the @ylecun and @lilianweng reposts of state-of-AI commentary are opinion without new evidence.