Summary
The one piece of real news is Qwen3.8-Max: a 2.4-trillion-parameter sparse mixture-of-experts flagship, generally available now with open weights promised next week alongside a 27B sibling. Hugging Face's Elie Bakouch supplied the sharpest read on it, noting that the two largest open-source models in the world now both use linear attention, which turns a research preference into a de-facto architectural standard for open frontier weights. Two smaller items are worth the click: Kilo Code published a measurement study of 10,643 real AI code reviews across 13 models and found open-weight reviewers match closed leaders on critical findings, plus a third of teams already use a different model to review than to write. Cursor shipped a token-efficiency improvement to cloud agents, and Prof Tom Yeh posted a by-hand Switch Transformer walkthrough that is the best free explainer of why MoE models can be enormous to store yet cheap to run. Everything else is filler: roughly fifty of this slot's sixty-four tweets are US and French political commentary from accounts that carry no AI signal at all, and they can be skipped without loss.
Posts
- Qwen3.8-Max lands at 2.4T parameters, open weights next week (@eliebakouch · Qwen blog · model card). Alibaba's flagship MoE claims 10-day autonomous coding runs from empty folder to production, native visual understanding through plan-execute-verify, and a 27B open-weights companion. Bakouch's observation is the real signal: the two biggest open-source models now both run linear attention, which makes the hybrid-linear direction the default rather than the experiment. See attention mechanisms and the Ling/Ring 2.6 hybrid-linear write-up.
- Kilo Code measures 10,643 AI code reviews across 13 models (cluster of 3) (@kilocode · full study). Findings normalized per review so high-volume reviewers do not win by default: open-weight models matched closed leaders on critical findings, and the apparent security gap collapses once one outlier model is removed. The operational number is 32.3% of attributed reviews using a different model than the one that wrote the code, with Step 3.7 Flash writing and Laguna M.1 reviewing the most common pairing. Author-reviewer separation is becoming a deployment pattern, not a research idea. Relevant to agent benchmarks.
- Switch Transformer worked through by hand in 13 steps (@ProfTomYeh). A manual walkthrough of the 2022 Fedus, Zoph and Shazeer paper that made sparse mixture-of-experts practical, where each token routes to a single best expert instead of every parameter firing. Worth bookmarking as the explainer to hand anyone asking how GPT-4, Claude, DeepSeek-V3 and Kimi can be huge to store and still cheap to serve. Pairs with today's DraftExpert, which reuses MoE sparsity for self-speculative decoding.
- Cursor cloud agents get 20-30% more token efficient (@cursor_ai). The gain jumps to 80% on runs involving computer use, from better handling of MCP servers, skills and screen interaction. Token cost is the binding constraint on delegating long-horizon agent work, so efficiency releases matter more than capability announcements right now.
- Enterprise vibe-coding hits a data-residency wall (@Scobleizer). The argument: internal apps built in Lovable or Replit are only useful wired to company data, but those tools execute in vendor cloud, so security's only lever is blocking the connection entirely. A real governance problem, though the post resolves into a vendor pitch, so read it for the framing and not the conclusion.
- Google DeepMind Post-AGI Research team Q&A at Deep Indaba (@GoogleResearch). Alex Goldin on sociotechnical implications of AGI, plus recent work from the team. Event announcement with no substance attached, but the existence of a named Post-AGI Research Team is itself a small signal about how DeepMind is organizing.
- Grok Imagine ships character, prop and location references (cluster of 2) (@heavypulp). Lock a character image and voice, then summon it in any video prompt with an @-mention for consistent appearance across generations. Live on Grok Heavy and Grok Plus. Identity persistence is the practical blocker for video generation in real production work, so this is the feature that matters more than resolution bumps.
- San Francisco AI billboard opens with "worried that AI will take your job?" (@Scobleizer). An AI sales company using job dread as the hook. Amusing as a read on how the local market is positioning itself.
- AI worker gives up on the job, in video form (@stepango). A joke clip making the rounds. No claim attached.
- VITURE teases a product reveal for 2026-08-06 (@Scobleizer). Save-the-date post with no detail. Skip.
- Podcast appearance announcement (@ns123abc). Self-referential promo. Skip.
- Political commentary flooding the feed (cluster of ~50) (@MarioNawfal · @brivael · @spencerpratt · @HouseGOP · @AustinJustice). Iran negotiations, Gaza, Ceuta, NYC politics, French intellectual history, municipal budgets. Zero AI content beyond passing mentions of robots and Israeli AI ambitions. Skip.