Summary
The only topic that genuinely repeated across slots is the SpaceX and NVIDIA hardware thread: the morning carried Starmind AI1 satellites with Vera Rubin NVL72 payloads at 250 kW each and up to a million satellites contemplated, and the evening closed the loop with Musk committing SpaceX to NVIDIA exclusively and calling Vera Rubin the best architecture available. The day's strongest single item is a morning standout, Cursor open-sourcing Mixture-of-Kittens, a fused mixture-of-experts training kernel claiming 2.37x over the fastest public baseline and 1.41x in their own 512-GPU production training, and the underdiscussed property is that it is bitwise deterministic, which makes MoE training runs reproducible in a way they currently are not. The evening's own standout is personnel rather than technical, with Demis Hassabis stepping back from running Google DeepMind and Jeff Dean leaving after 27 years to found Discovery Loop alongside Ghemawat, Le and Vinyals, announced four minutes apart, and Google funding the spinout rather than fighting it. The sharpest tension of the day sits across two slots: Anthropic confirmed an in-house inference-silicon team with Samsung in early talks on 2nm, on the same day the buyer building toward 10 GW committed publicly to merchant NVIDIA parts, so the custom-ASIC bet and the buy-NVIDIA bet are now explicit and opposed. Beneath those, a thin layer of small releases worth the click, Shieldstral, Alpamayo 2 Super, Qwen-Image-3.0 and Kiro Crew, plus tinygrad getting code execution on an AMD 7900XTX through custom firmware, which is the same instinct as Cursor's megakernel, taking scheduling away from the layer that normally owns it, arrived at independently on different hardware. Signal density was poor and should be stated plainly: of roughly 120 tweets across the two slots, about 70 were US, French and Middle East political commentary from accounts that are only nominally AI-adjacent, and the curated @bayesiansapien repost feed was empty in both slots.
Posts
- Cursor open-sources Mixture-of-Kittens, a deterministic MoE training megakernel for NVL72 (cluster of 3) (@cursor_ai · blog · GitHub) [morning]. Mixture-of-experts training spends most of its time moving tokens between GPUs rather than computing, and Cursor measured the MoE layer at over half of end-to-end training time, so MoK fuses all of that communication and computation into one persistent kernel. Claimed 2.37x on an isolated MXFP8 forward pass at expert-parallel degree 64 and 1,070.2 versus 760.9 tokens per second per GPU on 512 real GPUs, with HuggingFace's Elie Bakouch calling it "insanely good" and flagging the missing nvfp4 path. See the full write-up.
- Hassabis steps back from running Google DeepMind; Jeff Dean leaves to found Discovery Loop (cluster of 2) (@ns123abc · second post) [evening]. Hassabis becomes Chair of Google DeepMind and Chief Scientist of Alphabet, refocusing on AGI and disease through Isomorphic Labs, with the quote "I feel AGI is close at hand." Four minutes later Dean's exit after 27 years, taking Sanjay Ghemawat, Quoc Le and Oriol Vinyals into a public benefit corporation that aims to automate the experimental loop in ML and science, with Google investing and supplying compute.
- Anthropic confirms an in-house custom silicon team, with Samsung in early talks on 2nm (cluster of 3) (@ns123abc · Business Insider) [evening]. On the record via a job listing and a company statement: silicon engineers at $320k to $485k to co-design hardware and models together for Claude, targeting an inference chip, with Samsung already a Series H investor. Read against today's digest finding that Google structured roughly $200B of Anthropic chip lease contracts off its own balance sheet, and against the same vertical-integration move OpenAI made with Broadcom.
- SpaceX and NVIDIA put NVL72 racks in orbit, then Musk commits SpaceX to NVIDIA exclusively (cluster of 3) (@nvidia · @ns123abc · evening post) [morning + evening]. Starmind AI1 satellites carry Vera Rubin NVL72 payloads with Rubin GPUs and Vera CPUs at 250 kW peak each, with a stated ambition of up to a million satellites and 100 GW of orbital compute. No launch cadence, thermal or downlink numbers accompanied any of it, and those decide whether this is a product or a press release. Relates to the wiki's Vera Rubin gigascale ramp coverage.
- SpaceX Q2 earnings call carries the Grok roadmap and the Colossus power curve (@ns123abc) [morning]. Grok 4.6 lands roughly next week, 4.7 in three or four weeks, Grok 5 before year end, and the falsifiable bet is training on "the entire corpus of SpaceX data" from a quarter century to make it the best engineering model. Over 2 GW online by end of 2026 and "closer to 10 GW than 5 GW" by end of next year, set against $16B burned on $7.8B of revenue with $18.4B of capex.
- ByteDance bans employees from distilling US frontier models (@ns123abc) [evening]. The stated reasoning is political rather than technical, that distilling US models would draw Washington's attention and put TikTok at risk, so short-term gains get sacrificed. If true this is the first case of distillation policy set by regulatory exposure rather than capability economics, worth adding to the knowledge distillation picture.
- AWS open-sources Kiro Crew for running scheduled agent crews (cluster of 2) (@mattsgarman · blog) [morning + evening]. A workspace that takes a ticket queue, triages, dispatches and identifies owners, with agents wired into existing tools and schedulable for overnight migrations, on-call rotations and cross-repo incidents. Worth reading against today's finding that agent skill libraries mostly fail to abstract from experience, since scheduling more agents does not fix agents that do not learn.
- tinygrad got code execution on an AMD 7900XTX via custom MEC firmware (@tinygrad) [morning]. Custom firmware for the micro engine compute block, which handles command processing on AMD GPUs, ran its first kernel. The stated trajectory is that after spec and HCQ2 "the next phase of tinygrad will be operating system," a claim they intend to displace the vendor driver stack rather than sit on it.
- Diffusion Transformer worked by hand in 14 steps (@ProfTomYeh) [evening]. A manual walkthrough of the architecture that replaces the U-Net in a diffusion model with a transformer, using one training video and the prompt "sora is sky" at diffusion step t equals 3. The framing is the useful part: OpenAI shut Sora down this year, but DiT is what every current video generation model is built on, so the brand died and the mechanism did not.
- Mistral ships Shieldstral, a 3B open-weights content-safety model for on-device use (@MistralAI · announcement) [morning]. A small open-weight classifier meant to run locally rather than as a hosted moderation API, which matters because the default forces every moderation call through a round trip to the same vendor whose model you are moderating. No benchmark numbers or policy taxonomy in the captured content.
- NVIDIA's Alpamayo 2 Super goes commercially licensed for robotaxis (cluster of 2) (@nvidia · NVIDIA blog · weights) [morning]. Pitched as NVIDIA's most capable open reasoning model for autonomous driving, aimed at long-tail rare events rather than everyday scenarios, with inspectable decisions. Inspectability is the interesting claim, since the regulatory case for autonomy turns more on explaining a decision than on aggregate accuracy.
- Qwen-Image-3.0 lands on Qwen Cloud, second overall in the text-to-image Arena (@Scobleizer · model page) [morning]. Prompts up to 4.5k tokens, dense layouts with images inside images so newspapers and exam papers generate in one pass, text legible to 10px, 12 languages natively. The positioning is the giveaway, that it targets usefulness as "a deployable productivity tool" rather than good looks, which is a bid for document and UI generation rather than art.
- A $53k desktop runs DeepSeek V4 Flash on two Blackwells (@tinygrad · original) [morning]. Elliot Arledge reported running DeepSeek V4 Flash 0731 with dspark on two RTX Pro 6000 Blackwells, and tinygrad turned it into a hardware pitch with room for two more GPUs later. The datapoint that matters is not the price, it is that a frontier-tier coding model now fits on two workstation cards.
- Anaconda acquires EnkryptAI, folding AI risk detection into the platform that owns Kilo Code (@kilocode · @anacondainc) [morning]. A platform consolidating a governance layer under itself, sold as security and governance being the missing piece for enterprises scaling agentic development. Note the timing against SkillJack, which showed poisoned experiences laundered into reusable skills evade detection at 11.4% versus 98.5% on the source trajectory, exactly the risk class this covers and exactly what current scanners miss.
- NVIDIA argues enterprise AI needs a system of models, not one model (@nvidia) [morning]. Vendor content whose conclusion is convenient for a company selling the infrastructure to run many models, but it is also the industry restatement of the routing thesis this wiki tracks. It arrived the same day as a paper arguing the uncertainty signal every mixture-of-adapters router uses is the wrong signal.
- Kimi launches what it calls the world's first AI-native credit card (@hexiang · kimi.com/aicard) [morning]. Spending earns model tokens instead of air miles, with Select and Max tiers whose benefits are entirely quota: agent quota, Kimi Code at up to 20x, parallel multi-agent execution, scheduled tasks. Issued with Agricultural Bank of China, and bundling inference quota into consumer credit is a customer-acquisition channel nobody in the US has tried.
- Pika opens an API club at $10 per month across 100+ generative media models (@minchoi · dev.pika.art) [morning]. Positioned bluntly as "Gen Media is overpriced, we're fixing it," claiming up to 88% cheaper than other aggregators, with Seedance 2.0 R2V 480p at roughly $0.0452/sec against Fal.ai's $0.084 and Runway's $0.36. The interesting part is the spread between aggregators for identical models, which suggests margin currently sits in the aggregation layer rather than the model layer.
- Shopify posts $116B in merchant sales, up 32% year over year (@dhh) [evening]. $3.6B revenue up 34%, $654M free cash flow at an 18% margin, fifth straight quarter of GMV growth above 30%. A rare case of AI-adjacent growth showing up in cash flow rather than in capex commitments, though the company's claim that every shift AI creates plays to its strengths is the part to actually test.
- SF cannot hire AI and robotics talent while Seattle engineers absorb Amazon layoffs (@Scobleizer) [morning]. Every company Scoble spoke to in the last month names hiring as its biggest blocker, across engineers, researchers, hardware engineers and electricians, while Seattle software engineers struggle after Amazon cut thousands. Anecdotal, but it is the only labour-market signal in the day's feed, and the split is by specialization rather than geography.
- Stacked pull requests are obsolete now that you can delegate to an agent (@JonasBadalic) [morning]. An xAI engineer arguing the benefit of stacked-PR workflows died the moment the work could be handed to an agent. An assertion rather than an argument, but a specific claim about a developer-tooling category being made obsolete rather than improved, from inside a frontier lab.
- Scoble on cognitive outsourcing (@Scobleizer) [morning]. "For the next two years using AI will make you stupider. Then they will sell you cognitive enhancements," with a report that Musk said Neuralink is working on it. Included as an unusually candid statement of the dependency thesis from inside the SF scene rather than from a critic outside it, not because evidence is attached.
- Apple briefly pulled Telegram from the App Store over one moderation evasion (@dhh) [morning]. Durov's account is that a user planted illegal content in a public chat to trigger automated enforcement and take the app down, with restoration within hours. No AI content, but using an automated moderation system as a weapon against its own host is structurally the same laundering-through-an-automated-step problem SkillJack describes today in a different domain.
- Grumbling about Claude quality and about scraping law (cluster of 2) (@ns123abc · second post) [evening]. One post asking why Opus feels like it is getting worse over time, one asking when OpenAI and Anthropic will be punished under the Computer Fraud and Abuse Act. No evidence attached to either, but the first is worth tracking as sentiment if it keeps recurring.
- Omarchy Quattro ships three bespoke C++/Qt apps at about 1MB combined (@dhh) [morning]. Omacut, Omawrite and Omacalc, theme-synced, in roughly a megabyte total. No AI angle, noted only as a size datapoint in a week otherwise measured in gigawatts.
- Wispr Flow ships Notetaker (@Scobleizer · product) [evening]. Meeting notes with per-speaker attribution, pitched as clean enough to feed straight into agents and MCP. Skip.
- xAI's Imagine promotes world generation (@imagine) [morning]. "Build entire worlds based on your imagination" attached to a user video, with no capability claim, benchmark or technical detail. Skip.
- Google Research booth programming at Deep Learning Indaba (cluster of 3) (@GoogleResearch) [evening]. An AI trivia quiz, a WeatherNext tropical cyclone forecasting demo, and a Q and A on the sociotechnical implications of AGI. Event logistics only. Skip.
- Roughly seventy tweets of political and lifestyle commentary carried no AI content (cluster of 70) (@MarioNawfal · @brivael · @spencerpratt · @NICKIMINAJ · @WHFraudTF · @HouseGOP · @AustinJustice · @DoWCTO · @heavypulp) [morning + evening]. Iran shipping lanes, Ukraine drone strikes, French domestic politics, cartel indictments, LA housing, celebrity disputes, and a Dow record whose only AI content is the phrase "strong AI earnings." Two brushed AI adjacently, Trump telling communities that refusing data centers is a mistake and Flock's 80,000 cameras generating 20 billion plate reads a month with under 1% touching any crime, but neither carried enough to stand alone. Skip.