Summary
The morning belongs to one document, and it is not a paper. A DeepSeek kernel engineer named Shengyu Liu published a long personal essay overnight arguing that AI will write better CUDA kernels than he can within six to twelve months, and that he is going to keep accelerating it anyway on distributional grounds. It is a cluster of six posts spanning English and Chinese accounts, and two of them carry screenshots of the original text rather than a summary, which makes this the rare social story where the primary source is actually in the feed. Around it sits the second half of the pacing-the-frontier argument: Sayash Kapoor and Arvind Narayanan dropped a 13,000-word essay rejecting the framing on both sides, Gary Marcus amplified a funding-flow critique of the safety nonprofit layer, and a Google DeepMind alignment researcher announced his resignation. The research signal is thinner but sharper than the discourse: a Princeton architecture proposal called RLT that feeds each token's final hidden state back into the next token's input, a self-distillation method that improves reasoning with no teacher and no labels, and a Tencent sparse-attention paper being amplified a day after it was ingested here. Three bookmarks landed, all pointing at papers this wiki already holds, which is worth noting on its own. And the genuinely new commercial signal is that the agent harness has become organizational infrastructure: Y Combinator open-sourced the multi-agent harness it runs its own company on, and MCP standardized lazy skill loading in the same window.
Posts
The DeepSeek kernel engineer's essay, and it is the day's dominant story (cluster of 6: @AndrewCurran_, @Thom_Wolf retweeting @teortaxesTex, @MaxForAI, @hsu_steve, @CrazyShyyt, @ns123abc). Shengyu Liu, whose code is in DeepSeek V4.1's main attention kernel, describes a capability shift with a date on it. A year ago AI could help him look up documentation, read code and hunt bugs. Now it reads CUDA, PTX and SASS directly, profiles stalls per instruction with professional tools, and iterates on the optimization. PTX is NVIDIA's intermediate assembly and SASS is the actual machine code the hardware runs, and that level is where the last stretch of kernel performance lives, which is also where public documentation is thinnest. His estimate is six to twelve months until AI-written kernels match or surpass his own. The labour claim is more specific than the usual one: he does not expect unemployment, he expects the job to change from writing kernels to piloting an agent that writes them, and he says industry demand has already drifted from "people who can write high-performance kernels" to "people who can use AI to produce high-performance kernels faster." What he names as the loss is the craft, the afternoons spent tuning performance like a speedrunner chasing his own record. See the essay summary and the gpu-kernels page, where this is the first practitioner testimony against five benchmark results produced by people trying to prove the same thing.
The two screenshots in @MaxForAI's post are the original text, and they carry a nuance every English summary dropped (@MaxForAI). The first image is a section headed 那我呢, "What about me then?" It reads: when the day comes that AI writes kernels better than me, my judgment is that I will not quite be unemployed, but I will have to change professions. My rice bowl can still be preserved, but this may mean I never again get the chance to do the work I once loved. He then recalls an older judgment about himself, that because the era changes so fast he cannot predict five or ten years out, but that his vision, judgment, initiative and intelligence would keep him at the table. And he concludes that the judgment only guarantees he will not be unemployed, not that he will not have to switch fields, and that it in fact encourages him to avoid unemployment by switching. The second image is the closing section, 结语, which contains the two-extremes argument: either productivity is liberated and living standards rise, or a few tech companies control most resources while most people run weak models, and class mobility becomes a dead loop because you need the strongest AI to gain resources and resources to access the strongest AI. The parenthetical nobody translated is the most telling line in the essay: after describing the good outcome he writes, in effect, "I'll just write this much, otherwise I'm afraid it won't pass review." A censorship aside, inside an essay about who gets to own intelligence.
The Hitler comparison is real and it was explicitly flagged as hyperbole by its author (@ns123abc). The post headlines it as "Deepseek engineer compares Anthropic controlling AGI to Hitler obtaining the atomic bomb." The attached screenshot, a machine translation of the same conclusion, shows the full sentence: "I don't want Anthropic to master the most advanced artificial intelligence or AGI. It is an exaggeration to say that it is as serious as letting Hitler master the atomic bomb technology before the Allies." The qualifier is in the source and the headline drops it. The surrounding paragraph is the part that actually matters and it is the essay's thesis: cutting-edge intelligence should be supplied to everyone openly and cheaply, he does not trust Anthropic or OpenAI to do that, and this is why he stays at DeepSeek, to research powerful, fast and inclusive AI and open-source it. Worth reading the screenshot rather than the tweet.
Kapoor and Narayanan publish 13,000 words on what pacing the frontier should actually mean (@sayashk). Their most substantial safety writing since AI as Normal Technology, built on a month of analysing loss-of-control incidents at AI companies. Nine arguments, and the shape is a deliberate refusal of both existing camps. They agree with security practitioners that OpenAI did not adequately control its agents, while rejecting the implication that this is just applying thirty-year-old security methods to a new domain, because AI control is genuinely not a solved problem. They argue that marginal investment in control will pay off more than marginal investment in alignment, since the incidents show control being under-emphasised despite known techniques existing. They reject treating rogue agents as inherently catastrophic in favour of naming specific risks, with cyberoffense as the urgent one because agents can carry it out autonomously. The recommendation aimed squarely at Amodei is the fifth: organizational governance should be the primary tool, and AI companies are trying to reinvent it as a technology problem, when a single misconfigured RL environment can cause real harm and no individual team should be able to run that experiment without legal and security oversight. They also update in public, conceding that their earlier essay under-weighted risks arising during development and evaluation rather than deployment, and underplayed jaggedness. The full argument is in the pacing debate summary.
A Princeton PhD candidate proposes RLT, a transformer that feeds the previous token's final hidden state into the next token's input (@code_hiyouga, repo). Yifan Zhang, previously with ByteDance Seed and NVIDIA's Nemotron team, calls it the Recurrent Looped Transformer. A standard transformer already reaches earlier information through attention and the KV cache, which is the stored key and value tensors that save recomputing past tokens. RLT adds a direct feedback channel on top, so the next token's computation builds on the previous token's final representation and intermediate information can travel without ever being written out as text. The example configuration is a 48-layer encoder and a 48-layer decoder, and unrolling the recurrent state across 100 tokens gives a path through 4,800 decoder-block evaluations while reusing the same parameters, which is what the "infinite depth" framing means. The post is unusually honest about the cost, and this is why it is worth reading rather than skipping: within a sequence, decoder updates must happen in token order, so even when the whole prompt is available those updates cannot be computed in parallel, and a fixed number of blocks per token does not guarantee faster processing. The repository has small synthetic state-tracking experiments and nothing at scale. This wiki already holds a page on looped transformers and ingested the RLT proposal on 09-13.
Negative Self-Distillation improves reasoning with no teacher model and no ground-truth labels (@weizhepei, paper, code). Distillation normally means a strong teacher model generates outputs and a smaller student learns to copy them, which requires owning or paying for the teacher. NSD removes both the teacher and the labels: the model improves by learning to avoid self-constructed negative conditions, meaning deliberately induced reasoning flaws that it then learns to steer away from. Reported gains are up to 7.5% average across seven reasoning benchmarks, held across 1.7B, 4B and 8B models, which is the size sweep that makes a self-improvement claim credible rather than a small-model artifact. The cost angle is direct: every teacher-based distillation result in this wiki prices in teacher inference, and this one does not have that line item.
Tencent's SAS sparse-attention paper gets amplified a day after ingest (@CrazyShyyt). The post's framing is overheated but its technical description is accurate. Sparse attention puts a small learned selector in front of attention that scores blocks of past context and keeps only the best few, which is how you serve a million-token window without paying quadratic cost. The problem is that picking the top few is a discrete operation, so no gradient flows back into the selector, and every trainable method has worked around that by training the selector to imitate what full dense attention would have weighted. SAS argues that target is wrong under a fixed budget, because what you want is not "which blocks did dense attention like" but "which blocks, if I can afford only K, most change the prediction." It injects the selector's continuous scores directly into the attention softmax so the ordinary language-modeling loss trains it. The result the post gets right is that the gains are largest exactly where the compute budget is tightest. Ingested here yesterday as SAS.
Anthropic's agent-memory architecture circulates as a five-layer stack with a token-cost claim attached (@choopyplug1). The layers as described: working memory (the context window, where most agents stop), episodic memory (timestamped interaction logs, so the agent recalls that the deploy failed Tuesday because the migration script had a typo), semantic memory (facts and relationships as a knowledge graph, where "user prefers TypeScript" lives and does not expire with the session), procedural memory (the approach that worked becomes a reusable skill), and forgetting (an agent that never forgets accumulates contradictions, recommending restaurants in the city the user moved away from). The numbers quoted are secondhand from the PDF: Mem0 storing 1,800 tokens per query instead of 26,000, and Snowflake adding one ontology layer for 20% better accuracy and 39% fewer tool calls. Treat the figures as unverified, but the forgetting layer is the one worth sitting with, because it is the only one that removes rather than adds and the wiki's agent memory page has consistently found deletion to be the unsolved half.
An Anthropic and Edinburgh result claims unconstrained test-time compute actively degrades accuracy (@marfinxx). The four named failure modes are worth recording even though the post is secondhand and heavily embellished: distractor fixation, where irrelevant context turns a trivial counting task into a hallucinated multi-step algorithm; template forcing, where the model recognises a superficial pattern and imposes a complex formula on a problem needing arithmetic; spurious correlation drift, where the model abandons a strong empirical prior for background noise as token count grows; and deductive paralysis, where unguided deliberation causes second-guessing of valid intermediate steps, with a reported drop from 70% to 30% accuracy on dense constraint tracking. This matters today because it may be the mechanism behind the day's strongest paper. Today's HuggingFace Elo-per-token analysis finds agents convert tokens into quality faster than independent sampling early and then fall below it, and if the marginal late token is more likely spent second-guessing a valid step than finding a better one, that is exactly the curve you would observe. The poster's own production numbers (41.2% to 88.6% completion, 68% less wasted compute) are self-reported and should be discounted entirely.
A routing arbitrage claim with striking numbers and no paper attached (@slash1sol). The post describes a paper titled "The End of Model Loyalty" arguing that staying on one flagship is a tax rather than a preference. The setup: two frontier models now share the same million-token window, one renting at $10 in and $50 out and one downloading at $3 in and $15 out with open weights, with overlapping benchmarks on the median task and a gap only on the hard tail, quoted as 99.9% against 62.7% completion. The conclusion is a router rather than a favourite: price every job before dispatch, send the bulk to the cheap model and reserve the ceiling for finalists, with a reported production split of 84% to the cheap model and 16% to the ceiling, costing $256 for a workload the flagship alone would price at $4,318. No arXiv identifier or link accompanied the post, so this is a hypothesis and not a result. It is recorded because the shape is exactly what the llm-routing page has predicted since per-token pricing was shown to rank models backwards on task cost, and because a ~17x saving would be the largest production routing figure in the wiki if a real paper carries it.
Y Combinator open-sourced the multi-agent harness it runs its own company on (@stretchcloud). QM is MIT licensed and self-hostable, used internally across accounting, legal, events and engineering, giving each person and room its own scoped memory, files, credentials, permissions, crons, web apps and sandbox. The load-bearing design choice is that the underlying model is swappable between Pi, OpenCode, Codex and Claude Code without rebuilding workflows, which is vendor-neutral orchestration at organizational scale. The poster's read is that the harness is becoming what the database was in the 2000s: every serious org will have one, most will not build it from scratch, and the ones who understand the architecture will compound an advantage. What makes it notable rather than another launch is that it is the first prominent open-source harness built by an organization for its own operations and then released, and it ships permissions and sandboxes while deliberately not shipping domain verification. That is the harness page's portability boundary confirmed by a release decision: structure travels, evidence does not.
MCP now defines an extension for discovering and loading agent skills from MCP servers (@dani_avila7, repo). The flow is deliberately plain: the server advertises available skills, the client fetches skill metadata, and the
SKILL.mdbody loads only when actually needed. That is lazy delivery standardized at the protocol layer, and it is the industrial form of a research finding this wiki already holds, that a small always-on core plus a per-task retrieved subset beats injecting one comprehensive playbook on every step, at a fraction of the tokens. What the protocol does not supply is the selection policy, which remains the open question.A curated stack of ten agent skills with 3.49M combined downloads (@beamnxw). Named repos include mattpocock/skills, vercel-labs/agent-skills and supabase/agent-skills, spanning planning, React performance, databases, interface systems, search, browser work and production UI. Useful mostly as a distribution measurement: skills are now shipped by framework vendors as a first-class artifact rather than assembled per project, which is the precondition for the MCP extension above mattering.
A textbook recommendation that is genuinely worth the click (@burkov, book). Andriy Burkov points at Foundations of Large Language Models by Tong Xiao and Jingbo Zhu, 277 pages and 76 illustrations, covering the full lifecycle from pre-training data and architecture through alignment, inference acceleration and system-level scaling. The last two are the reason it is listed here rather than skipped: a textbook treatment of inference acceleration is rare and this reader's core areas are exactly the chapters most books omit. Free to read with an account.
AlphaXiv's OpenResearch was the top trending GitHub repo on Friday (@askalphaxiv, repo). It turns a coding agent into a research agent that reviews literature, forms hypotheses, runs experiments and produces artifacts, now with Windows support. Relevant here because alphaxiv's paper overviews already feed this wiki's deep dives, and because it is the same closed loop that today's Dream-RSI and Atria Dawn papers formalize, shipped as a repository instead.
A Google DeepMind AGI safety researcher announced his resignation (retweeted by @Miles_Brundage). Bilal Chughtai, who worked on AGI safety and alignment research at Google DeepMind, posted his departure. No detail beyond the announcement in the captured text, but the timing inside the pacing-the-frontier week is why it is circulating.
Gary Marcus amplifies a funding-flow critique of the AI safety nonprofit layer (@GaryMarcus, original by @kevinnbass). The argument is that safety nonprofits orbiting Anthropic, from Tarbell Fellows in media to METR in evaluations, are ultimately funded by Dustin Moskovitz's roughly $7B Anthropic stock held by the parent funding organization, which appreciates as Anthropic grows, complicating independence claims. It is a conflict-of-interest argument rather than an evidentiary one, and it is worth holding alongside Amodei's proposal that embedded third-party evaluators are the fix, since the question of who pays the evaluator is exactly the unresolved part.
An ex-OpenAI researcher's case for what he calls sane AI regulation (@oleg_murk). Five years at OpenAI, and the framing is that you do not need to believe in doom to fear the next decade or in utopia to be excited. The one concrete detail in the captured text is the claim that OpenAI recently used roughly 10,000 concurrent AI agents on a single task, which if accurate is a useful scale marker for the agent-swarm argument running through today's digest.
A multi-agent permission pattern stated more cleanly than most papers manage (@gippp69). The observation, attributed to OpenAI's agent team: eight agents can think, research, reconcile, draft and verify, but only one gets permission to commit. Intelligence scales out while authority stays narrow, through an intake, extract, reconcile, draft, verify, controlled-write pipeline. Unsourced, but it is the same blast-radius principle the harness literature has converged on and it compresses well.
Three saved posts landed, and all three point at papers this wiki already holds. They are covered properly in today's Media Zone rather than here.
Skip. An AI engineering roadmap graphic, a small-language-model monetization thesis, a headless-browser launch, a Marvin Minsky multi-agent framing post, and the Amanda Askell pile-on. No research substance.