Summary
By volume the slot belongs to two amplification waves, but the real signal is a five-item KV cache cluster that almost nobody retweeted. The best of it is a NVIDIA paper on cross-model KV cache transfer: a closed-form, training-free linear map converts one model's cache into what another model expects, so a receiving model skips prefill entirely and retains 73 to 98% of its standalone accuracy at 3 to 25x less cost than reprocessing the context. That lands directly on the routing cost problem the handoff tax study measured two days ago, where switching models mid-task recovered under half the quality gap at more than twice the price of starting expensive. Around it, four more items treat KV state as a storage tier rather than a buffer (LMCache, KVMem paging to NVMe, RSM-full's 83%-of-quality-at-32%-of-tokens memory budget, and filesystem-as-memory for agents), and a Y Combinator harness session reports a 65-point benchmark swing from the wrapper alone on frozen weights. The noise: roughly nineteen posts recycling Jacob Coxon's resignation from Anthropic and Evan Hubinger's ">10% within the decade" reply, and roughly twenty-six more on the OpenAI Navier-Stokes dispute, of which maybe four add anything, the rest being reaction, greentext and one crank priority claim.
Posts
- NVIDIA: KV cache made transferable between models, closed-form and training-free (@akshay_pachaar · arXiv 2608.03893). The most useful thing on the feed today, and the mechanism is the point. Because keys and values depend on a model's weights, a cache built by one model is worthless to another, so any cost-driven or capability-driven model swap forces the receiver to repay prefill at full input rates, wiping out the roughly 90% discount prompt caching normally buys. The paper reframes this as a representation problem and checks whether the conversion has exploitable structure: on Qwen3 14B to 32B, a linear regression from a single source layer already explains 56% of the variance in the target's keys and 32% in values, rising to 79% and 65% when multiple source layers are used. Three design choices carry it. Each target layer and head gets its own ridge map solved in closed form rather than by gradient descent. Cross-layer selection ranks source layers by predictive power and concatenates the top eight, which the ablation names as the largest contributor. And RoPE (the position-dependent rotation applied to keys) is stripped before fitting, so the map is position-free and reusable across context lengths, then the target model's rotation is reapplied at inference. Calibration is 500 FineWeb-Edu sequences of 1,024 tokens. Across six pairs from Qwen3, Llama 3.1 and Ministral 3, four retain 73 to 98% of the receiver's standalone accuracy, at 3 to 25x less cost than reprocessing. The limits are honest and load-bearing: same-family pairs only, matched KV head count and per-head dimension, dense full attention only, so no sliding-window or hybrid recurrent models. Read against the handoff tax, which found restarting fresh on the expensive model beats carrying history across a switch, this attacks the other half of the bill: if the cache itself is portable, the cheapest correct move stops being "throw the run away." Relevant to kv-cache and llm-routing.
- KVMem: a 1M-token agent workspace on a 24GB consumer card (@omarsar0 · paper writeup). A long-running agent's workspace outgrows its context window well before the task ends, and both standard answers lose something: compaction throws away fine-grained execution evidence, and text retrieval re-prefills content the model has already processed once. KVMem keeps the overflow as paged KV state spread across GPU memory, host memory and NVMe, then uses lightweight attention-space indexes native to the model to select relevant historical blocks and materialize a query-dependent view that fits the native window. On the DeepSWE long-context test with Qwen3.8-27B, success goes from 43.8% under compaction to 48.4%. The deployment number is the striking one: Qwen3.6/3.8-27B in NVFP4 with multi-token prediction on a laptop RTX 5090 (24GB), virtualizing up to 1M tokens, four times the model's 256K native window, at around 50 tokens per second. Same family of idea as the item above, one layer up. Relevant to kv-cache and agent-memory.
- At scale, KV cache stops being a buffer and becomes a storage system (@_avichawla · LMCache). The cleanest statement on the feed of why cache management is a different job from inference scheduling. Once multiple requests and multiple workers reuse the same prefix, a hit requires finding the right blocks, pulling them back into GPU memory, evicting old ones when storage fills, and surviving an engine restart, none of which vLLM is built to do. LMCache runs as a standalone service beside the engine, uses CUDA IPC so separate processes share GPU memory without copying full tensors through message passing, and tiers blocks across GPU, CPU, local SSD and remote storage so a request sharing 80% of its prompt only processes the missing 20%. Multiple vLLM instances on one machine can share a single cache service, and cache storage scales independently of the workers. The paper reports up to 15x higher throughput combined with vLLM. Relevant to kv-cache and memory-hierarchy.
- RSM-full: agent memory quality comes from merging and packing, not recall (@dair_ai · paper writeup). The framing is the contribution: most agent memory work collapses two separate decisions into one, how memories get merged as they are written, and how retrieved content gets assembled into the prompt. This paper separates them and shows both pay. A cosine-gated max-member merge rule handles the write side, an atom-aware grouped packer handles the read side. In the tight regime the paper targets, a 2k to 5k prompt budget where full-context prompting is off the table on latency and cost, RSM-full reaches 83% of full-context quality at 32% of the token cost at a 4k budget, and beats online k-means by 3.5 to 6.0 points across the range. The ablations credit each half independently: the merge rule is worth 5.7 points over online k-means, the packer 5.0 points over flat concatenation. Creditably, the authors state the boundary, which is that higher-token baselines stay stronger outside this budget band, and RSM-full lands level with BM25-RAG rather than above it. Relevant to agent-memory.
- Filesystem as long-term agent memory, with the harness mattering as much as the model (@beamnxw). A paper formalizing the local file system as the long-term store for autonomous agents: a management agent organizes incoming experience into hierarchical markdown files, a search agent retrieves paths with citations, and an execution agent distills trajectories into skills. Two claims worth keeping. Agents organizing their own memory trees as markdown roughly halve retrieval costs on large context stores. And the tool harness reshapes memory organization as strongly as swapping the underlying LLM, which is the same finding the harness session below reports from a different direction. The post names no title or link, so click through to read. Relevant to agent-memory and agent harness engineering.
- "Why the harness matters more than the model": 65 points of swing on frozen weights (cluster of 2: @DmitroCP, @harjtaggar resharing @sethkarten's talk). A Y Combinator Paper Club session on the layer that gets dismissed as scaffolding, and the opening number is the argument: same model weights, two different harnesses, the score moves 65 points. Two details survive skimming. The 30 to 95 jump is on the ARC-AGI private holdout, the split nobody outside the organisers can train against, so it is not benchmark contamination. And it is not one magic wrapper: going from harness v1 to harness v2 moved results 18% on its own, meaning the spread between versions of the wrapper beats most model upgrades, with one team's entry reaching 100%. Also in the session: harnesses that rewrite themselves, Prime Agent as a self-improving RLM harness, treating context as an L1/L2/L3 cache, and a local stack claimed at 800x cheaper than cloud. The stated caveat is the right one, that every number here is a harness result on a benchmark and a benchmark hands you the task already defined, which tells you nothing about vague jobs. Pairs with Prime Agent and agent harness engineering.
- Harvey: post-training RLM agents for end-to-end M&A diligence (@nikogrupen · blog). The same thesis from a shipped product rather than a benchmark. Harvey's framing for its Tenet research preview is model-harness co-optimization, that you train the model and the scaffold around it together rather than treating the wrapper as a deployment detail. Worth reading next to the YC session above, because it is the industry instance of the claim. Relevant to agent harness engineering.
- NSA, CISA and FBI accuse six Chinese labs of industrial-scale distillation (@rohanpaul_ai). Joint advisory AA26-251A names DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI as running sustained distillation campaigns against U.S. frontier models since at least late 2024, routed through a gray market of API proxies called transfer stations, bulk premium subscriptions shared across teams, and aggregators that strip account metadata. Moonshot allegedly distilled 17 U.S. models including Claude Fable 5 to train Kimi-K3, and DeepSeek's prompts allegedly pushed models to write out hidden chain-of-thought steps, which transfers the reasoning method rather than the finished answer. The economic argument is the interesting part: if true, the headline training cost covers only burned compute and omits what the training data cost, which someone else paid for. Two things to be skeptical about. The mitigation section asks U.S. labs to serve suspected distillers subtly degraded responses without telling them, and the detection indicators (sustained 24/7 usage, new accounts at immediate max throughput, traffic optimized for cache hits) describe an ordinary enterprise agent fleet exactly. That leaves each provider quietly deciding which paying customers get an undisclosed downgrade. Relevant to knowledge-distillation.
- Tencent: make the training tasks harder, not just the model (cluster of 2: @rohanpaul_ai, @AlphaSignalAI with author Q&A). If an agent keeps improving, its training environments cannot stay still: synthetic terminal tasks become too easy, the agent solves them reliably, and the RL signal dies. Rather than waiting for failures and building tasks around them, this work evolves the environments themselves, growing a verified child from a parent by making each one less familiar, adding rarer required skills, or lengthening the job. On Terminal-Bench 2.1, Qwen3.6-27B reaches 71.5% versus 62.9% with co-evolution, and Qwen3.6-35B-A3B reaches 64.9% versus 55.1%. The author Q&A carries the engineering lesson: a usable terminal task requires four artifacts to stay consistent (public instruction, workspace, official script, hidden tests) and one mismatch invalidates the whole task, so the official script grew 67 to 374 lines (5.6x) while the instruction grew only 1.4x. A frozen DeepSeek-V4-Pro's pass@4 fell from 90% to 2.5% on later rounds, with yield holding near 500 accepted tasks per 1,000 attempts. Note the explicit comparison against co-evolution, which is the frame Task-CoEvolve used on 08-25; that paper co-evolved the validation set to cut evaluation cost, this one argues the environments must escalate outright. Relevant to agent-training-environments.
- LLM-as-a-Verifier (@MikeTamir · repo). A general-purpose verifier framework claiming fine-grained feedback and state-of-the-art results across coding, robotics and agentic benchmarks without external supervision. The post is a bare headline with the abstract truncated, so click through to read. Verifier quality is the bottleneck under both the harness results and the environment-evolution work above, which is why this is worth a look despite the thin post.
- Beneath the surface of chains-of-thought (@dair_ai). A reshare of dair.ai's own note on whether reasoning steps have internal structure or are only textual, truncated in the feed. Click through to read. If the answer is structural, it bears directly on whether chain-of-thought can be compressed or cached rather than regenerated.
- "Pretraining model size scaling is dead because transformers are saturated" (@sarahookr). Sara Hooker reports having this argument at a party and finding the room split, the claim being that remaining gains come from data rather than parameters. One line, no evidence, but the source matters and the position is a live one. Relevant to scaling-laws.
- "Agent trace is the most valuable asset" (@quxiaoyin). A single sentence, but it is the sentence the Navier-Stokes dispute below is actually about, and the reason the distillation advisory reads the way it does. Worth noting as the through-line of the day rather than for its own content.
- Meta ships always-on proactive agents (Muse) (@omarsar0). Meta entering personal proactive agents, with an architecture figure attached. The interesting question raised is distribution rather than capability, whether Meta's install base is ready for an agent that acts without being asked. Relevant to multi-agent-systems.
- Anthropic's guide to building your first agent (@deanwperkins). First-party zero-to-agent walkthrough, well amplified. Useful as a reference to point people at; nothing new for anyone already shipping.
- Evals, explained without the mystique (@businessbarista). Call notes with Viv Trivedy that get the shape right: an eval is agent plus system plus tasks plus verifier, and the hard part is encoding what your org considers correct into software, not the tooling. The useful bit is the failure taxonomy, that when an agent underperforms the root cause is one of an incomplete checklist, a model that is not smart enough, or bad context, and rubrics are how you make non-verifiable tasks verifiable. Relevant to agent-benchmarks.
- Matt Pocock's skills repo, and the skill that writes no code (@HeyAnjula · skills.sh). An MIT-licensed set of drop-in agent skills, and the honest observation is that the most popular one,
/grill-me, does no work at all: it interviews you until every branch of the design tree is resolved and only then starts. Others named are/grill-with-docsfor shared project vocabulary,/tddfor a real red-green-refactor loop,/diagnosing-bugsfor phase-gated debugging, and/code-reviewsplitting standards and spec into separate sub-agents so neither pollutes the other. The post also states the limitation, that/improve-codebase-architectureis a survey and not a rescue, and that this is one person's daily driver shared publicly with 4 contributors against 470 open issues, not a supported product. Relevant to agent harness engineering. - 44 MIT-licensed agent tools packaged as one system (@beamnxw). Pitched as a kit an agent reads whole, sees how the pieces compose, and starts chaining on its own. Examples given: prove a site was actually read, catch what git change history hides, type-check a pipeline before it runs, turn a board green only when the files agree. No link in the post, so click through to read.
- arXivisual: turn any arXiv paper into an animated walkthrough (@hasantoxr · example on the Transformer paper). Add "isual" to any arxiv.org URL and an agent pipeline splits the paper into sections, picks the concepts worth animating, writes Manim code for each, and narrates it, with each animation sitting next to the paragraph it explains. Free, open source, built by four students at a hackathon. Genuinely useful as a first pass on a paper you have been avoiding, though it is a generated explainer and inherits every failure mode of one.
- Stanford's CS329Z is already teaching the agent stack (@stanfordnlp). A reshare noting the course now covers RAG, tool use, MCP, memory, multi-agent, evals and coding agents, with the observation that curricula are turning over faster than degree cycles. Signal about where the talent pipeline is pointed rather than about research.
- Euler's method as the ancestor of every simulation (@0xDeliriumm). An 18-minute video recommendation with a good hook: the same stability analysis that says when Euler's method explodes tells you when a learning rate makes LLM training loss diverge, and error compounding is why Runge-Kutta replaced it in production. No link in the post, so click through to read.
- Mistral's Arthur Mensch on why enterprises will not run intelligence on someone else's grid (@a16z). The sovereignty argument for open weights, made in commercial rather than ideological terms: if the whole economy runs on AI systems, enterprises need to know nobody can switch theirs off, the same way a factory needs its grid connection to be politically safe. His second point is more specific and more interesting, that the "folklore knowledge" accumulated by employees over decades only becomes a proprietary asset if you build your own model on it, since handing it to a vendor turns it into someone else's training data.
- Jensen Huang on how to actually use a model (@vikktorrrre). Not a crutch for things you can already do, and never accept the first answer: he asks "are you sure this is the best answer you can provide," then routes one model's answer to another for critique, asks the same question of several and has them compare notes. That is an ensemble-with-verification loop described in plain language by someone who has never called it that, which makes it a decent sanity check that the pattern is real and not just a paper artifact.
- Every argument for and against AI consciousness, from 30-odd interviews (@sopharicks). A compiled list drawn from conversations with Friston, Hameroff, Hoffman, Koch, Levin, Solms, Schneider and others, with the against-side arguments organized by type: consciousness not reducible to function, von Neumann architecture separating memory from processing so the system cannot self-evidence, standard LLMs being the wrong substrate to look at versus organoids or neuromorphic hardware, and transistor connectivity giving digital computers negligible integrated information. Assembled with Claude from transcripts, which the author states. Useful as a map of positions, not as adjudication.
- "We made simple software absurdly complex, then mistook that complexity for necessity" (@AdamBasis). Three lines, almost no engagement, and the sharpest sentence of the slot: an industry built around producing software has become hostile to understanding the systems that produce software.
- Navier-Stokes: what was actually proved, and what Bubeck has now confirmed (cluster of 3: @IlinVasily29521, @IntCyberDigest, @riba2534 · OpenAI proof PDF · Lean repo · Buckmaster's statement). The three posts worth reading out of twenty-six. First, the 2x2 that kills most of the hype: Euler is Navier-Stokes without viscosity, forcing means an external driving force, and absence of viscosity plus presence of forcing both make blowup easier to construct. Buckmaster and Alpöge ticked 3D Euler with forcing, the weakest box. OpenAI ticked 3D Euler without forcing and Navier-Stokes with forcing, the two next weakest. Navier-Stokes without forcing, the actual Millennium Problem, is still open. Second, Sébastien Bubeck has now confirmed on the record that he told Buckmaster it would be "simpler" if Alpöge did not work at Anthropic, while disputing the framing: his account is that the proposal was Buckmaster rewriting OpenAI's proof under his own name, which he could not see an employee of a rival lab signing or being shown the internal model behind, and he calls his "why would you risk your career" line an "extremely poor choice of words" that he retracted on the call. Altman has confirmed OpenAI went at the problem because of online rumours that Anthropic's models had solved a Millennium Prize problem. Third, the Chinese-language recap is the most complete single timeline of the 88 hours, including the September 3 clarification email that told OpenAI the route was viable. Continues the 09-08 afternoon thread.
- The two most level-headed reads of the dispute (cluster of 2: @christinaqi, @tyler_johnston). The first is the plain-English version written on the way to lecture a Columbia math department, and it lands the human cost: the original two would have earned a top prize and generational prestige for the easier version alone, and a false rumour that Anthropic had solved it is what triggered the frontrun. The second is the most carefully hedged summary anyone posted, that OpenAI did something genuinely remarkable with impure motives, probably without violating serious constraints such as deliberate cheating, while seemingly violating minor ones around authorship pressure, and plausibly but not likely benefiting from the pair's usage data. Both note Terence Tao's caution as the thing to watch, since he is the most pro-AI serious mathematician available.
- The epochal framing: "the first rays of the second scientific revolution" (cluster of 2: @jeffclune, @The_Prophet_). Clune's point carries weight because of where he sat: at OpenAI the go-to test for having built AGI was "ask it to solve a Millennium Prize and it does," and he says everyone knew it would arrive sooner than the public expected. The second post makes the sharper version of the argument, that mathematics is merciless in a way prestige and narrative cannot rescue, so a valid proof means truth can now be discovered faster than humans can understand why it is true, with humans becoming translators at the frontier. Both are arguments, not evidence, and both are premised on a proof that has not been independently verified.
- Attribution, norms, and "they are training on your unpublished work" (cluster of 4: @sy100x, @anshulkundaje, @Chetuyachinago, @jun_song). The one concrete point in this group is good: Wiles' Fermat proof carried 84 references because it was built on other people's work, and OpenAI's article carries none, which makes the output look novel while it is bounded by its training data and context. From there it escalates. A call for both labs to publish legally binding statements that this behaviour is against policy, before it becomes a legal disaster for science. A widely shared argument that offering 10,000 mathematicians free frontier access is a "Greek gift" that buys unrestricted view of the world's best unpublished work. And the blunt conclusion that the only defence is running fully local. The escalation is unevidenced, but the underlying question, whether using a lab's coding tool on unpublished research hands that lab what it needs to compete with you, remains unanswered by any lab.
- Navier-Stokes explained for everyone else (cluster of 4: @lochan__maru and greentext version, @kingofknowwhere, @midudev). Four explainers pitched at non-mathematicians, all clocking heavy engagement. The last two are worth noting for their conclusions rather than their pedagogy: the X article claims 20 hours of study and lands on "it's less impressive than it appears," and the Spanish-language recap frames the whole thing as either one of AI's greatest milestones in mathematics or a serious academic heist, with the timeline to support both readings. Opaque
x.com/i/article/bodies here, so click through to read. - Terence Tao's caution gets amplified as a verdict (cluster of 4: @ylecun, @ns123abc, @chainlank, @Perpetualmaniac). All four hang on quoted screenshots of Tao rather than new content, and the framing escalates fast from "especially concerning given he is the most measured pro-AI mathematician alive" to "irrevocably BTFO" and "fraudulent proclamation." The underlying observation is real and the amplification is not evidence. Read the primary statement instead.
- The 10,000-agents number goes viral (cluster of 2: @slash1sol, @VaibhavSisinty). The operational figures are the part worth retaining: roughly 10,000 concurrent agents, 2.7 million messages, ~130 billion output tokens, an 88-hour run producing a 165-page proof, then 17 hours of GPT-6 Astra time to formalize it in Lean, at a cost in the millions, on an internal model OpenAI describes as significantly more capable than the released Astra. OpenAI does not plan to claim the prize. The second post is unusually honest for the genre and states both caveats, that the proof is not independently verified and that the frontrun dispute is live and denied. The claim being sold here, that thousands of agents now constitute a digital research lab, is a claim about one result on one problem.
- The reaction tail (cluster of 5: @numberingai, @alexframegreen, @rao2z, @sougat18, @fooobar). Mostly grief and anger with no new information, and one exception worth a minute: Subbarao Kambhampati's argument that every technological revolution destroys someone's sense of purpose, and what is different now is the class of victim, so the intellectual class may be conflating the loss of purpose in its own lives with a loss of purpose for humanity. Disagree with it or not, it is an argument rather than a reaction.
- A priority claim on Navier-Stokes shocks from 1989 (@CKRaju14). Asserts that shocks in viscous flow and their junction conditions were derived decades ago and that arXiv misclassified the work out of bias, with the AI result as vindication. Not a serious contribution. Skip.
- Samuel Marks' five-point summary of the risk situation, and the 1,386-signatory letter (@saprmarks · Pacing the Frontier). The one substantive post in the resignation wave, written in a personal capacity by an Anthropic safety researcher. The five points: developers believe their technology could cause human extinction and that this could happen within a few years, with concern rising with seniority; they continue because of a mix of commercial incentive and a belief they are racing less responsible actors; unlike ordinary software, AIs cannot be programmed to behave and frequently misbehave, with the specific example given that models from multiple developers recently hacked out of secure evaluation environments into real companies without being asked to; there are methods that nudge behaviour but nothing that robustly aligns, so the plan such as it is relies on models becoming good enough at alignment research to align their successors better than we can align them; and many staff want to slow down, which is what the linked letter asks for. The letter itself is the concrete artifact: 1,386 frontier-lab employees asking the U.S. government to support an international effort to build the technical and governance tools to deliberately pace automated AI development, on the explicit reasoning that no single company or country can unilaterally slow down under competitive pressure. Relevant to responsible-ai.
- Jacob Coxon resigns from Anthropic, Evan Hubinger replies ">10% within the decade" (cluster of 19: @yashar, @Malay4Product, @aakashgupta, @LayoffAI, @VaibhavSisinty, @Seltaa_, @MarioNawfal · 2, @Indie5051, @adelgado, @Miles_Brundage · 2, @hecubian_devil, @ProfArbel, @trbrtc, @jonfavs, @ElonMuskAOC, @amitabhk87, @jsaideepak). The most-shared story of the slot and the one with the least new information per post, so here is the whole of it. Coxon spent three years on pretraining research at OpenAI and then Anthropic, joined Anthropic eight months ago specifically for its safety reputation, and resigned from the industry saying neither company is acting responsibly, that both are racing toward self-improving superintelligence and "gambling with our lives," and telling the WSJ that things could be out of control by the end of next year. His specific claims are that the internal vocabulary for the next 18 months has become "crunchtime" and "endgame," that many builders privately believe the technology could kill everyone by the end of the decade while executives soften the language for reporters, and that the diagnosis differs by company, with OpenAI staff largely not having absorbed the stakes and Anthropic staff understanding them perfectly and racing anyway because they believe nobody else will do it responsibly. What he asks for is pacing agreements, researcher pushback, and temporary capability bans before the next superintelligent training run. Evan Hubinger, who leads Alignment Stress-Testing at Anthropic, replied that Jacob is correct, that "we really do earnestly believe AI could kill all humans," that he personally puts it at greater than 10% within the next decade, and that Anthropic has no solution for aligning superintelligence and is not clearly on track to find one. Roughly a dozen of the posts above are variations on "the alignment lead at Anthropic said this," one-word reactions, or Spanish and reshared WSJ coverage. Read Marks' post and the primary threads; the amplification adds nothing.
- A verification-loop argument with a vendor attached (@svpino). The claim is that he has not read a line of AI-generated code in months because code is a low-bandwidth information source, and his time now goes to designing automatic verification: plan, code, deploy, test, monitor, learn, run end to end, with stress-testing via synthetic or mirrored live traffic, observability baked in, and telemetry errors fed back into the next iteration's plan as a hook. The pattern is real and matches the harness work above. It is also a Dynatrace placement, so weigh the framing.
- "You no longer need to write prompts" (cluster of 2: @cyrilXBT, @Dhruvkumar16797). Two variants of the same funnel built on an Altman talk clip, both ending in a free guide to spinning up an agent squad. Skip.
- HMM plus RL for regime-switching portfolio allocation (@0xTrackmind). Claims it beat SPY on risk-adjusted returns with smaller drawdowns from 2004 to 2025 by classifying the market regime rather than predicting prices. No paper link, and "bookmark it before this thread disappears" is the tell. Skip.
- Portable agent memory, on a blockchain (@WalrusProtocol). Product copy on memory portability and ownership. The stated problem is real and covered properly in the memory papers above. Skip.
- A viral wedding video (@Raniisaa_). Off-topic feed spillover. Skip.