Summary
The evening feed is the pacing-the-frontier fight going geopolitical, and it added exactly one new fact: China's Foreign Ministry answered Dario Amodei officially, spokesperson Guo Jiakun calling the proposal "fearmongering, confrontation and vicious competition," with Trump and Xi now set to meet on September 24 to discuss AI governance. Fourteen of the thirty-eight posts are that story or a reaction to it, mostly re-narrated in four languages, so the cluster is loud and thin. The real technical signal is small but unusually good: a compressed MLSys optimization map covering Flash Attention, GQA, MLA, RadixAttention, speculative decoding, ZeRO and DualPipe in one thread, the full FT detail on Latham & Watkins buying its own Nvidia racks and fine-tuning Nemotron 3, a self-rewiring routing graph that strengthens and atrophies edges by outcome, and two releases from the same Princeton author, the Recurrent Looped Transformer and FlashREINFORCE. Sakana's PC-ALM, a local-learning alternative to backpropagation that trains thousand-layer networks, slipped through as a bare retweet and deserves more attention than the feed gave it. The rest is a Navier-Stokes crank post, a prompt-guide funnel and the usual recap threads.
Posts
The MLSys optimization map, compressed into one thread (@shao__meng · CS336 · Scaling Book). A Chinese-language walkthrough of Gauri Gupta's frontier-lab interview notes, and it is the single most useful item in the slot. The inference section is the densest: KV cache evolution from grouped-query attention to DeepSeek's multi-head latent attention, which compresses keys and values into a low-rank latent space and cuts the cache by an order of magnitude, then cross-layer KV sharing and Gemini-style interleaved local plus global attention. It names stateful caching as the serving-layer split that separates candidates who know models from candidates who know deployment, rolling hash plus tree-structured prefix cache plus LRU eviction, which is RadixAttention. Speculative decoding gets the right framing, 2 to 3x decode speedup with the output distribution mathematically preserved, which is why it beats naive distillation or pruning. Training side reads as a distributed-training decision tree: ZeRO's three sharding stages, Ring All-Reduce's 2×(N−1)×X/N per-card transfer cost, pipeline bubbles from GPipe's (d−1)/(m+d−1) through PipeDream weight stashing to Zero Bubble and DeepSeek-V3's DualPipe, Megatron column versus row tensor splits, context parallelism, and auxiliary-loss-free MoE balancing. The thread even flags its own error, noting bf16 has fp32's dynamic range so loss scaling is really an fp16 concern. → KV cache
Latham & Watkins, now with the model name and the headcount (@bearlyai). The afternoon slot carried this as a framing post. The FT detail is the part worth keeping: the firm will spend "potentially hundreds of millions of dollars" on in-house legal AI built by fine-tuning Nvidia's open-weight Nemotron 3, with about 100 of its 900 technology specialists on the project. The CIO gives two reasons and they are the two that matter for anyone modeling this trend, client information too sensitive for any cloud vendor, and "with a lot of consumption costs coming, we wanted flexibility." That is a customer explicitly pricing lab API consumption as the thing to escape, not just a privacy story. → compute economics
A routing graph that learns its own wiring (@0xCodio). Packaged as engagement bait with a "don't let this rot in your bookmarks" close, so discount the delivery, but the mechanism is a clean statement of routing as a learned policy rather than a fixed topology. Four operations: TRACK records outcomes per edge, meaning the specific handoff between two nodes rather than the nodes themselves; STRENGTHEN thickens an edge when work that flowed through it ended well; ATROPHY thins edges that keep producing bad outcomes until the graph stops using them; REWIRE runs continuously so the map is redrawn from results. The reported result is 12 edges withering to nothing and 5 new preferred routes emerging in a month, routes the author says he would not have designed. The claim under the story is that the best path through a system is not knowable in advance and only shows up in which paths keep paying off, which is the same argument the routing literature makes about static model-selection rules. → LLM routing
Recurrent Looped Transformer, explained against the Ouro and Astra line (@0xLogicrw). The clearest short explanation of RLT yet on the feed, and it draws the distinction most coverage misses. A standard transformer reads the past through attention and the KV cache. RLT additionally hands the internal state left over from computing one token directly to the next token, a relay baton rather than extra laps. Ouro and Astra style recurrence loops the same token through the same layers several more times, which is checking your answer again before submitting. RLT passes state between different tokens, so per-token depth is unchanged but the computation chain grows with sequence length, which the paper calls infinite time depth. → Recurrent Looped Transformer
FlashREINFORCE, from the same author as RLT (@yifanzhang_ · repo). Critic-free, single-rollout, asynchronous RL for agentic language models, released today alongside the RLT thread above by the same Princeton researcher. Critic-free and single-rollout is the cost claim: no value network to train and one generation per prompt instead of a group, which is a direct attack on the memory and sampling overhead that makes GRPO-style training expensive. Announcement only so far, no numbers in the post, so the repo is the thing to check. → RL for LLMs
The RSI roadmap, plus the sharpest objection to it (cluster of 2: @Gorden_Sun · paper, @TanayAyitmaz · Self-Adapting LMs). The Chinese summary walks the five levels cleanly: L1 executes a human-written improvement procedure and banks the artifacts, L2 diagnoses its own weaknesses and picks among improvement methods, L3 chooses or generates its own training material for the gaps it finds, L4 adapts in deployment and distills day-to-day work into a durable tool and skill library, L5 rewrites the research, diagnosis and self-training machinery itself and passes the stronger learning ability to the next version. The Turkish post is the useful counterweight and it is the more grounded read: a trillion-parameter model updating its own weights, middleware and harness is a myth today, and what actually works is small agents running an advanced hyperparameter search and data refinement loop, where batch size, gradient accumulation, learning rate, optimizer and the ZeRO or DDP choice already decide whether even a LoRA run succeeds. → Last AI Built by Humans
Sakana's PC-ALM, backprop-free training at 1000 layers (@hardmaru). A bare retweet with the text cut off, which is why it got almost no traction in this window, but the claim is the kind that matters if it holds: a local-learning alternative to backpropagation that trains thousand-layer networks using only local signals. Local learning removes the global backward pass, which is the thing that forces activation storage and serializes the layer dependency chain, so the memory and parallelism implications are the interesting part rather than the accuracy number. Click through to read the full announcement.
DeepSeek attributes real gains to automated environment construction (@adithya_s_k). DeepSeek is now naming automated RL environment construction as a post-training pipeline component and crediting a significant share of its performance improvements to it. That is a lab confirming what the agent-environment papers have been arguing all month, that generating the training environment is the bottleneck rather than the training algorithm. The post points at Repo2RLEnv as the closest open-source implementation of these recipes. → agent training environments
China answers officially, and the answer is no (cluster of 6: @choblin29, @Hesamation, @MarioNawfal, @angeldot_, @AnatoliKopadze, @VaibhavSisinty). The one new fact in the slot's dominant story. Foreign Ministry spokesperson Guo Jiakun, asked directly about Amodei's proposal, said "fearmongering, confrontation and vicious competition will only disrupt the process of global AI governance, which serves no one's interest," and called for AI to stay "open and inclusive." That is Beijing answering officially, where the earlier Cold War playbook line came from state-run Global Times. Hesamation's post carries the strongest version of the Chinese argument: Amodei says global pacing requires China's cooperation while simultaneously proposing restrictions on its chips, compute and models, names China twelve times in the essay, and writes that those restrictions "would slow China's progress enough to widen America's lead significantly over the next 3 to 5 years." Trump dismissed the proposal the same day with "whoever wins AI, wins," and Trump and Xi meet on September 24 on AI governance. Five of these six posts are the same wire story in Spanish, English and Persian, so treat the volume as amplification, not corroboration. → pacing the frontier
The regulatory-capture read gets its financial version (cluster of 4: @vincentweisser, @nate_taplin, @AndrewOrlowski, @ZackKorman). Taplin's is the substantive one because it lists the balance-sheet events rather than asserting motive: OpenAI cut prices 50% to fend off mostly Chinese open-weight models, its Q2 operating loss rose by about $3 billion or roughly 30%, the IPO was delayed on Altman's "ill-advised moment" language, data-center backlash is about to reshape Congress, and Z.ai and Moonshot shipped models rivaling Claude in a possible agentic DeepSeek moment. His conclusion is that a business with OpenAI's retail spending commitments may not close without restricting low-cost open-weight competition, and that being deemed "safe" by regulators is a way to cut R&D spend without losing share ahead of an IPO. Orlowski's jab is that METR, a proposed evaluator, did not notice when it was itself hacked. Korman's is the cleanest principled objection, that there is no singular danger, only dangers with tradeoffs, and no one holds a monopoly on safety.
LeCun reopens the GPT-2 file, and the rebuttal is better than the attack (@sleepy0x13). Yann LeCun named Dario Amodei directly: he said GPT-2 was too dangerous to open-source, we mocked them then and everyone should keep mocking them. The Chinese thread lays out the history accurately, GPT-2 at 1.5B parameters held back in favor of a 124M release over fake-news and spam fears, the staged-release policy that followed, Amodei as Research Director at the time, and the full release nine months later with OpenAI's own retrospective finding no strong evidence of large-scale misuse. Then it does the harder work of pushing back in both directions. LeCun is pinning an institutional decision on one person, and Amodei's 2026 argument is about a different class of system, since he explicitly says asking for a slowdown in 2023 would have been pointless and what changed his mind is agentic behavior, cyberattack incidents and early recursive self-improvement. The question the thread lands on is the right one: is AI safety a risk science that reality can recalibrate, or a system where every failed prediction becomes "good thing we were careful."
Microsoft publishes a Code of Conduct for its own frontier models (@mustafasuleyman · code of conduct). A first draft governing MAI models as they approach the frontier, framed around AI being subordinate and in service of people. It is the third lab-authored governance artifact this week, so the pattern worth tracking is labs racing to write their own rulebook while governments decline to write one.
India's objection: you are pacing the wrong thing (@Product_nation · 2023 letter). The best-argued dissent of the slot. The labs are converging on pacing development, slower capability gains judged by evaluators they invite in themselves, which is conveniently the lever that costs everyone outside the frontier labs the most. The sharper risk is the gap between a finished training run and a billion users, now measured in weeks, and that is where independent oversight belongs. It also insists oversight cannot rest on a pledge, since pledges are revocable and self-scored, and has to be enforceable in architecture through verifiable evaluation, auditable chains and accredited actors. Closing line is fair: three CEOs agreeing on a Saturday afternoon is not global governance, and most of the world's AI users live outside the countries drafting these rules.
Who gets to be the independent auditor (cluster of 2: @timnitGebru · NY Mag, @timnitGebru). On Stanford as a proposed independent evaluator, quoting the line that the membrane between academia and industry there is practically nonexistent and it can be hard to tell where the university ends and the businesses begin. The second is a retweet of the claim that the Coxon story sits on top of a financially entangled network. Positional rather than evidential, but the conflict-of-interest question is the one the evaluator proposal has not answered.
Marcus argues for US-China collaboration in The Economist (@GaryMarcus). Excerpt from a new essay proposing that the relationship will be adversarial in parts but should include cooperation, with medicine and cybersecurity as the plausible joint agenda. Paywalled, so click through to read.
Claude Code has 118 commands and most people use ten (@claudeskills101). A survey of the ones that get skipped, and it is a decent harness inventory even if the framing is listicle. The interesting ones are the control surfaces rather than the conveniences:
/effortto set how hard the model thinks,/goalto set a condition to work toward instead of a single instruction,/advisorto bring a second model in purely to critique the first,/sandboxto isolate what it can touch,--worktreefor isolated parallel runs and--max-budget-usdto cap spend. Those last two are the cost lever, since parallel agent work and budget ceilings are the difference between a loop you can leave running and one you cannot. → agent harness engineeringAlibaba open-sources OpenSandbox (@0xJokker). An isolated execution environment for agents, one sandbox per agent, covering code execution, web browsing and full desktop control, running locally under Docker and scaling with Kubernetes. Over 15k stars. The sandbox layer is the unglamorous prerequisite for letting coding and GUI agents do real work, and it is now commodity rather than something each team builds.
Wireless charging as a power side channel (@thesupermanmx). Cornell researchers show the Qi standard leaks what is on your screen. Processor current draw while rendering a page echoes back through the magnetic field into the charger, so monitoring the transmitter's current identifies which of the top 20 websites you are browsing with up to 95% accuracy, with no software access, no permissions and no data connection. The attack hardware is a cheap microcontroller that fits inside a public charging pad. The counterintuitive detail is the best one: it works best above 80% battery, because a nearly full battery stops absorbing bulk charge so the delivered power tracks the processor workload closely.
Seven free repos people are paying for (@ayush26291). Ollama for local model serving, Dify for visual LLM app pipelines, Firecrawl for turning sites into clean LLM-ready markdown, prompts.chat as the prompt library, AutoGPT, a Nous Research repo and one more. Nothing new for anyone already building, but the Ollama plus Dify plus Firecrawl trio is a reasonable no-API-bill starting stack.
Oracle's layoff email, verbatim (@Amanda_Goodall). Same-day termination notice with DocuSign severance paperwork and an urgent request for a personal email address before system access is cut. No numbers on scale attached, but it is a data point on how the infrastructure vendors are reshaping headcount while announcing compute buildouts.
The week-in-AI recap threads (cluster of 3: @VaibhavSisinty, @simplykashif, @JasonObermaier). The seven-day recap covers the Coxon resignation at over 100 million views, the alignment lead's greater-than-10% extinction estimate this decade, Amodei's essay, Altman agreeing and ruling out a 2026 IPO, Musk's "Dario is right," and Nadella welcoming deliberate pacing. The second restates Altman's two failure modes, loss of control and concentration of power. The third is a one-line nudge at European politicians. All three are compression of things already covered, useful only as a timeline check.
RL meets systematic alpha research (@RuujSs). An agent that learns what survives validation and searches toward it, with every rejection as signal and the research loop reframed as an environment. The framing matches the environment-construction thread above, but there is no artifact, benchmark or link attached, so it is enthusiasm rather than a result.
Off-topic and promo (cluster of 4). Skip. @dr_logvinovich claims OpenAI's reasoning models hit "the topological limit of the universe" on Navier-Stokes and that the Clay Institute is reviewing it, which is a crank post wrapped in real equations, @0xRicker a "you no longer need prompts" funnel for a GPT-6 Astra guide, @enjojoyy a note about bookmarks outnumbering likes, @docmilanfar a truncated retweet about what AI discourse is missing.