Summary
The strongest thing in this slot is a chart, not a claim: a Bloomberg price-per-million-tokens comparison in which DeepSeek V4 Flash is so cheap its bar is invisible next to Qwen3.8-Max, Kimi K3, GPT-5.6 Sol and Fable 5, which sits at roughly $47 per million output tokens. Robert Scoble reads it as a survival question ("who can survive the price squeeze?") and pairs it 45 minutes later with a capability datapoint pointing the same direction, a report that Qwen 3.8 Max went from 0% to 42% on a hard induction benchmark and now ranks second ahead of Fable 5 at a fraction of the price, so treat those two posts as one cluster about Chinese open-weight models closing the capability gap while holding a 10x price advantage. The best practitioner item is a reproducible recipe for running MiniMax H3 text-to-video with audio on a single NVIDIA DGX Spark, which required working around SM121 compatibility, online FP8 quantization and a pinned vLLM-Omni build, and generates a clip in 2.5 minutes. Two carryovers from the overnight window are worth the click: Kilo's 10,643-run code review study, already written up as a routing page, and a 13-step hand-worked Switch Transformer walkthrough that is the cleanest available explanation of why frontier mixture-of-experts models are huge to store and cheap to run. Everything else is thin. Roughly four fifths of the 53 tweets are US, French and Middle East politics from three accounts that are AI-adjacent in name only, the curated repost feed is empty for a sixth consecutive slot, and the remaining AI items are wearables, mocap and Grok promotion.
Posts
Qwen and DeepSeek are squeezing frontier pricing from below (cluster of 2) (@Scobleizer quoting @AndrewCurran_, and @Scobleizer quoting @DeryaTR_). The attached Bloomberg chart plots price per million tokens, input and output, for five models, and the shape is the story: DeepSeek V4 Flash is at or near zero on both axes and does not render a visible bar, Alibaba's Qwen3.8-Max sits around $4 output, Moonshot's Kimi K3 around $12, OpenAI's GPT-5.6 Sol around $28, and Anthropic's Fable 5 around $47 output on roughly $9 input. That is close to a 50x spread between the cheapest and the most expensive, on models that are increasingly being compared on the same benchmarks. Curran's own reaction was that he first assumed Bloomberg had forgotten to add DeepSeek to the chart. The second post supplies the capability half: Qwen 3.8 Max reportedly jumped from 0% to 42% on a hard induction benchmark, moving into second place ahead of Fable 5. Read against SemiAnalysis on Kimi K3 serving, which showed a B300 node holding only 3.25M tokens of KV budget after weights and prefix cache hit rate collapsing below 10% above concurrency 8, the pricing spread is not purely a margin decision. Someone is either eating serving cost or has a materially better serving story, and this chart does not distinguish the two. The load-bearing open question is whether $4 per million output tokens is sustainable unit economics or market share purchased at a loss.
MiniMax H3 text-to-video running on one DGX Spark (@Scobleizer quoting @aijoey, repo). The repo documents a measured compatibility recipe for MiniMax H3 on a single NVIDIA DGX Spark, and the three things that had to be solved are the interesting part: SM121 compatibility for the Spark's compute capability, online FP8 quantization rather than a pre-quantized checkpoint, and a pinned vLLM-Omni build because the multimodal serving path is not stable across versions. The full text-to-video-with-audio workflow is now reproducible, with what worked, what broke and the exact smoke tests written up. A clip takes 2.5 minutes to generate, which the author concedes leaves work to do. This is the most useful kind of practitioner signal for anyone tracking GPU kernel and serving work: not a benchmark number but the specific list of things that break when you take a frontier multimodal model off a datacenter node and onto one desktop-class box.
Kilo's 10,643-run study on open-weight code review (@kilocode, research post). Kilo classified findings across 13 models in real code review runs and reports three results: open-weight models matched closed leaders on critical findings, the apparent security gap mostly disappears once a single outlier model is removed, and models disagree substantially on what severity deserves escalation. The routing detail is the one that matters most, that 32.3% of attributed reviews used a different model than the one that wrote the code. Full writeup at Kilo open-weight code review routing, and it sits directly under the standing question on the LLM routing page about whether authoring and reviewing should be separate model choices in production. Kilo's data says teams are already answering yes.
Switch Transformer worked by hand, 13 steps (@ProfTomYeh). A step-by-step manual walkthrough of the Switch Transformer (Fedus, Zoph and Shazeer, 2022), the paper that made sparse mixture-of-experts practical at scale by routing each token to exactly one best expert instead of several. Yeh's framing is the useful one for a newcomer: this is why GPT-4, Claude, DeepSeek-V3 and Kimi can pack enormous parameter counts while activating only a small slice per token, huge to store and cheap to run. Good background for the LatentMoE analysis in the Kimi K3 primer, where the same design choices Yeh explains conceptually turn out to be set by interconnect bandwidth rather than by anything about representation.
OpenAI publishes a point-by-point response to Apple's lawsuit (@Scobleizer quoting @OpenAINewsroom, statement). OpenAI calls the suit baseless, disputes Apple's claims about its employees, and says it is publishing messages documenting what happened. Two of the largest companies in consumer computing litigating over talent movement in public, with receipts attached, is a real escalation from the usual quiet settlement pattern. Click through to read the statement itself.
Anthropic's pay-versus-mission problem, and the joke version of it (cluster of 2) (@brivael quoting @LeaBourseFR, and @brivael quoting @0xDevShah). The first reports Anthropic's CEO saying he worries new hires are joining for the salary rather than the mission, and points out the obvious tension: Anthropic set that market itself, paying up to $900k for some profiles to win talent against OpenAI and Google, while reportedly targeting an IPO near a $1T valuation, so employees are watching their stock. The second is the compressed version of the same critique aimed at the safety posture, "anthropic: if we release this humanity might end / qwen: here is god in a zip file please star our repo," which lands harder than it should given the pricing chart at the top of this slot. Both are commentary rather than reporting, but the salary and valuation figures are the kind of number worth logging.
Goodfire on tearing an LLM apart to see how it thinks (@Scobleizer). Scoble reshares his own note on a lunch with Goodfire founder Eric Ho and Curt Tigges, framing the company as interpretability research aimed at building human trust in AI systems. No technical claim is in the post, but the commercial bet is on the right side of the wiki's current evidence: the filler-token result showed chain-of-thought monitoring has a demonstrated hole while hidden-state readouts held up, which is exactly the layer Goodfire sells.
Grok as a video forensics tool (@MarioNawfal quoting @elonmusk, shared conversation). The pitch is that you can upload any clip, have Grok go through it frame by frame, ask questions about specific details, and have it flag whether the footage may be AI-generated. An AI-detection feature shipped by a lab that also ships a video generator is a genuine conflict of interest worth watching, and no accuracy numbers accompany the claim.
Comic 4.2 estimates foot physics from plain video (@Scobleizer quoting @JonathanJarvis). An updated flagship perception model whose headline feature is foot physics estimation, recovering force, friction, toe roll and elevation change from ordinary video input with no LiDAR, usable directly in Unreal Engine with Maya and Blender support coming. Markerless mocap from a single camera is a narrow win, and the reason it appears here is the same reason the DGX Spark item does: another capability that used to need dedicated hardware now runs off commodity sensor input.
Wearable AI that still cannot name the person you are talking to (cluster of 2) (@Scobleizer, @Scobleizer). Scoble reports the Looki pendant camera produces a daily journal that he wishes were more useful, and names the missing feature outright: getting the names of the people he is talking to, which he acknowledges is freaky. He also notes pendant cameras make poor cameras because they are never aimed well, and that both he and his lunch companion forgot the device was recording. That last observation is the actual signal in an otherwise slight post.
Grok Imagine promotion (cluster of 3) (@MarioNawfal on turning a 2011 dog meme into a video, game or song in a day, @MarioNawfal on Optimus figuring out new situations rather than following mapped-out moves, and @Scobleizer resharing @sophiamyang putting AI-generated video inside a finger frame). Enthusiasm with no measurement attached in all three. Skip.
XMoney one-liner (@stepango). "XMoney is the new checkbook," in reply to an unrelated joke. Skip.
Politics and off-topic, grouped. The bulk of the slot. @brivael posts 18 items of French domestic political commentary covering Macron, Attali, Fabius pensions, air conditioning policy and subscriber revenue. @MarioNawfal posts 13 items including Saudi Aramco's 44% quarterly profit jump amid Strait of Hormuz disruption, Snoop Dogg threatening to sue Spotify over $45,000 for a billion streams, the Beirut port explosion anniversary, Ukrainian drone strikes on Russian logistics hubs, EU monitoring of Spain over Ceuta, and a Yunnan cliffside elevator. @spencerpratt posts five items on California wildfire policy, the Paramount-Warner merger challenge and US left politics, one of which attaches an LA Times report that the share of homes surviving wildfires is falling despite home-hardening efforts. None of it is AI content. Skip.