social-stream · 2026-08-13

2026-08-13

Summary

Only two slots ran today, morning and afternoon, across 173 captured tweets, and the day has exactly one dominant cluster: SpaceXAI's Grok 4.6, which surfaced across eight handles in both windows. The story worth keeping is the efficiency profile rather than the benchmark score, namely 61 on the Artificial Analysis Intelligence Index, roughly 53 agent steps where Claude Opus 5 needs 103, $2 in and $6 out per million tokens, and Elie Bakouch's model-card read showing the large gains sit on internal benchmarks SpaceXAI built itself (GPU kernel generation, optimizing its own chat inference) while the public software-engineering evals trail. The single strongest artifact of the day is a screenshot in the morning slot, the Terminal-Bench 3.0 leaderboard, which publishes tokens and dollars per run next to accuracy and names the agent harness on every row, confirming Bakouch's read from the outside and showing that token spend does not track accuracy at all. The afternoon's standout is hardware economics rather than models: Jensen Huang arguing that CUDA is what makes a 2020 A100 a rentable and financeable asset through 2029, which is a depreciation-schedule argument landing two days after NVIDIA's $500B financing platform, plus a Bloomberg report that Anthropic is in talks to buy Decart, a world-model and inference-optimization startup, for roughly $6 billion. Outside those four items the signal is thin, mostly unverified model-shop needling from one handle and a robotics pile that sits outside this wiki's range. There were zero @bayesiansapien retweets for a fourth consecutive day, so the curated layer is absent again, and roughly 105 of the 173 tweets were US and French domestic politics from a handful of handles with no AI content.

Posts

  • Grok 4.6 ships at Grok 4.5's price, and the efficiency claim is the real one (cluster of 8, @zhu_hanqing666 · @eliebakouch · @aksheyd · @ellev3n11 · @ns123abc · @minchoi · @dhh · model card) [morning + afternoon]. 61 on the Artificial Analysis Intelligence Index, five points over Grok 4.5 a month later, completing agentic workflows in about 53 steps against Opus 5's 103 at $2 in and $6 out per million tokens. In an agent loop every step is a full round trip carrying accumulated context, so halving steps compounds well past the sticker price. Full treatment in the Grok 4.6 summary.

  • Bakouch's model-card read: strong on the benchmarks they built, behind on the ones they did not (@eliebakouch · model card) [morning]. Big gains on DeepSearchQA and internal KernelBench, best-in-class on SpaceXAI's own engineer benchmark, state of the art on "inferenceEval" which measures optimization of their own chat inference, and behind the frontier on Terminal-Bench 3.0, SWE Marathon and DeepSWE. A legible profile, and unverifiable from outside.

  • "Grok 4.6 optimized the Grok Build harness for itself" (@aksheyd) [morning]. One sentence, and the sharpest line in either slot, landing the same morning HuggingFace published a paper showing a strong model writing a harness for a weaker one nearly doubles its accuracy without touching weights. It means nobody outside SpaceXAI can separate the step reduction into weights versus scaffold. See today's digest.

  • Terminal-Bench 3.0 leaderboard, read off the attached screenshot (@aksheyd, relaying @ryan_marten) [morning]. Rare in publishing tokens and dollars per run next to accuracy and naming the harness on every row: Opus 5 with mini-SWE-agent leads at 42.7% for $5.8k, Grok 4.6 with Grok Build is fourth at 26.5% for $2.1k, which is $79 per accuracy point against Opus 5's $136. Claude Code appears four times spanning 4.6% to 34.1%, and Sonnet 5 burns 17.9B tokens and $6.9k to land at 14.6% where Fable 5 in the same harness reaches 34.1% on 3.6B, so token spend is not tracking accuracy.

  • Jensen Huang on GPU fungibility: a 2020 A100 stays productive to 2029 (cluster of 2, @JensenHuang · @ns123abc · Business Insider) [afternoon]. The chain is that CUDA gives one platform across Ampere, Hopper and Blackwell, versatility makes compute fungible, fungibility drives utilization, and utilization makes GPUs rentable, durable and financeable. It is a depreciation-schedule argument dressed as a software argument, arriving two days after NVIDIA turned AI compute into an asset class, which needs exactly this claim to hold.

  • Anthropic in talks to acquire Decart for roughly $6B (@Scobleizer, relaying @shiringhaffary · Bloomberg). [afternoon] Decart is a world-model and inference-optimization company, and the second half is the part that matters here. If the price is real, a frontier lab is valuing a serving-efficiency team at world-model multiples, the clearest market signal yet that inference cost is a moat rather than a line item.

  • tinygrad and comma ship "chestnut," a $249 USB-to-PCIe eGPU dock (@tinygrad · comma.ai/shop/chestnut) [morning]. ASM2464PD-based, USB2/3/4 plus PCIe 4.0 x4, open-source firmware, $249 alone or $799 with a GPU. The real announcement is in the product page: the first chestnut-class driving model has 30x more parameters and 100x more FLOPs than comma's current on-device model, so comma is moving its on-vehicle compute tier up two orders of magnitude.

  • Google Research: recall, not storage, is the factuality bottleneck (@GoogleResearch · post) [morning]. A behavioral framework measuring encoding and recall separately finds that when a frontier model gets a fact wrong, it usually learned the fact and cannot retrieve it. That points at retrieval-side interventions rather than more pretraining data, and it rhymes with the day's tool-use result where retrieval returns do not affect the answer.

  • Bakouch's sovereignty argument about Mistral (@eliebakouch · Mistral post) [morning]. Mistral announced in-region inference, open-model serving and new European compute. His worry is not that Mistral's models trail the frontier but that governments and large companies may be forced to use them in security-critical roles like cybersecurity, so hosting the best available open models on sovereign infrastructure usefully decouples where compute runs from which model you must use.

  • Musk says Grok 4.7 lands in three to four weeks (@ns123abc · @minchoi) [morning]. Initial training is done and it is in supplemental training on "a massive amount of SpaceX company data." Read next to the model card's internal-engineering strength this is a consistent strategy, and it raises the same question the public-eval gap already raises.

  • DeepSeek quietly shipped V4 Pro, one point above V4 Flash (cluster of 2, @ns123abc · follow-up) [afternoon]. His claim is that there was no announcement anywhere because the gap is embarrassing, and the follow-up is him defending it against open-model boosters. Unverified and snide, but a one-point Pro-over-Flash delta is checkable against the V4 architecture page.

  • Ambrosia Energy: solar plus battery at $100/MWh with a claimed 12 months to power (@Scobleizer, relaying @FutureJurvetson · TechCrunch · ambrosia.energy) [afternoon]. Two SpaceX alumni pitching off-grid continuous renewable power to datacenters, with a 12x faster proprietary solar install as the actual differentiator. If contract-to-power really lands at 12 months, the binding constraint shifts from interconnect queues to install throughput, which is the cost-per-megawatt layer of inference as energy-to-token production.

  • SSI test-time-training rumor, flagged as fake news by the person relaying it (@ns123abc) [afternoon]. A small reasoning engine allegedly competing with much larger training runs through meta-learning-oriented data curation, plus gradient descent replacing the context window. Vapor, but the direction is what the continual-learning market is already pricing.

  • The "new" Reuters DeepMind recursive-self-improvement story is April reporting (@ns123abc) [afternoon]. A Brin-and-CTO-led strike team turning coding models into full AI researchers, trained on Google's private codebase. The debunk is the useful part, since the article's first line dates it to April.

  • DHH ships Omarchy Quattro RC, then Quattro plus the Aether theme builder (cluster of 4, @dhh · @dhh · PR) [morning + afternoon]. Cost optimization in miniature: another 114MB off the ISO by compressing the NVIDIA 580xx drivers at maximum zstd, three kernels plus all drivers plus the distribution under 6GB, a 56-second install on a $699 Dell XPS 13, and agents shipped as first-class through mise on an out-of-band upgrade path. His broader claim is that Linux adoption among programmers is about to go parabolic because the open-source advantage with agents is unstoppable.

  • X Square's WALL-B hits 1,816 parcels per hour at 98% accuracy (@Scobleizer, relaying @TheHumanoidHub) [afternoon]. Roughly two seconds per parcel in a fully autonomous logistics livestream, about 45% faster than Figure's sub-three-second pace from May. Scoble's own caveat is right: task-specific policy, not a generalized world model, and not a humanoid.

  • A clean taxonomy of why robots fail outside simulation (@Scobleizer, relaying @muskan_kalra24) [afternoon]. Three named gaps: sim-to-real physics, the sensorimotor gap where robots get much weaker pressure, texture and friction feedback than humans, and action representation. Outside this wiki's range, but unusually crisp for a thread.

  • Quintar pitches spatial AI as infrastructure for live events (cluster of 4, @Scobleizer · quintar.ai) [morning]. A 3D basketball court rendering that tracks everything in a stadium to within a few inches, with the hard problem framed as anchoring AI to the right real-world object at distance in a crowd rather than the AI itself. Recorded, not pursued, though the claim that grounding is an infrastructure problem rather than a model problem echoes the harness cluster.

  • "Stop trying to control AI, we are growing these things" (@Scobleizer, relaying @cxgonzalez) [afternoon]. The argument is that coercive framings in safety work will backfire as models get more lifelike, with a prediction that alignment researchers will rederive parenting from first principles. Pure vibes with no concrete claim, so it stays out of responsible-ai.

  • China's industrial capacity as the real war variable, and the yuan at a 3.5-year high (cluster of 2, @MarioNawfal · yuan) [afternoon]. One thread argues a prolonged conflict favors China on manufacturing capacity, rare-earth control and supply-chain depth, the other notes the yuan at its strongest since February 2023 despite tariffs and export bans. Tangential, but rare earths are the semiconductor story's upstream.

  • Lex Fridman's Khabib episode, dubbed with human plus AI translation (cluster of 2, @lexfridman · Russian version) [afternoon]. Recorded entirely in Russian, shipped with both audio tracks and subtitles. Only the production pipeline is on-topic, and only barely.

  • Promo pile: Tesla FSD anecdote, a Grok-built browser game, Design Arena usage, Tesla brand loyalty (cluster of 4, @Tesla · @stepango · @MarioNawfal · Tesla loyalty) [morning + afternoon]. A frame-by-frame claim that FSD Supervised began evading 0.17 seconds before the at-fault car moved with no fleet statistics behind it, a survivors-like game generated by Grok 4.5, and two agency-written posts with credited designers. Skip.

  • Off-topic political and personal volume (cluster of roughly 105: @brivael, @MarioNawfal, @spencerpratt, @WHFraudTF, @HouseGOP, @DoWCTO, @JonasBadalic, @SeanParnellASW) [morning + afternoon]. US and French domestic politics, Iran and Hormuz coverage, military news, local elections, Netflix documentaries, a Paramount merger lawsuit, and two tweets about moka pots. Skip, with one line worth keeping: a Nawfal thread relaying that roughly 40% of the world's helium, which chip manufacturing depends on, moves through the Strait of Hormuz and cannot be piped around it.