social-stream · 2026-08-04

2026-08-04-morning

Summary

The strongest signal in this slot is a speed claim with an explicit architectural qualifier attached: Celeris says its Celeris-1 model now holds the top spot on Artificial Analysis for response time, time-to-first-token and time per intelligence task, at 2,038 output tokens per second, roughly twice the next fastest of 591 models, while beating Grok 4.3 on Humanity's Last Exam, and the sentence that matters is that they claim it "without custom silicon." Second is a capacity statement from AWS CEO Matt Garman on Bloomberg Tech, that much of AWS capacity is already spoken for through 2027 and into 2028 with demand still significantly outstripping supply, which is the supply-side counterpart to every serving-cost result in today's digest. The slot's only genuinely new product surface is Cursor giving coding agents direct read and write access across Google Workspace, meaning Gmail, Drive, Calendar, Docs and Sheets, which is five more mutable external states for an agent to track in the same week two papers established that tracking mutable state is what agents do worst. There is one funding item, HUMAIN's first investment in a Saudi company, taking a stake in enterprise AI vendor MOZN to co-build sovereign AI for financial institutions and the public sector. Robert Scoble carries three of the slot's remaining substantive posts, all meeting notes rather than claims: Goodfire on interpretability, mixedbread on a database built for AI agents, and a 5,000 square foot private Holodeck in Austin that John Carmack and Meta's CTO have both flown out to see. There are zero curated reposts for the fifth consecutive slot and zero linked papers across all 70 tweets, and no article body in the farmer's capture is a paper or blog post worth a wiki page. Roughly four fifths of the feed is US and French geopolitics from two accounts that are only nominally AI-adjacent, so signal density stays at the low level recorded all week. The one thing the slot supplies that the research sources cannot is the market's tone: a record $75 billion into global tech equity funds over five weeks, posted without any accompanying argument about what it is buying.

Posts

  • Celeris-1 claims the fastest-model spot, and says it did it on commodity hardware (@Scobleizer quoting @Celeris_ai). The claim is that Celeris-1 is ranked #1 fastest of all 591 models tracked by Artificial Analysis, leading on response time, time-to-first-token and time per intelligence task simultaneously, at 2,038 output tokens per second, which they state is 2x the tokens per second of the next fastest model. They also claim it beats Grok 4.3 on Humanity's Last Exam, a benchmark of expert-written questions designed to resist saturation, so the pitch is speed without the usual capability sacrifice. The load-bearing clause is the last one: "Celeris-1 proves it's possible to deliver useful intelligence 10x faster than today's systems, without custom silicon." That distinguishes it from the inference-speed claims this wiki has logged from Groq and Cerebras, which are silicon arguments, and puts the gain in software and serving architecture instead. Nothing in the post says what the architecture is, and 2,038 tokens per second on a model that is competitive on Humanity's Last Exam is a strong enough combination that it needs independent confirmation of the model size and the batching regime before it means much. Worth reading against today's SemiAnalysis analysis of Kimi K3 serving, which shows that headline throughput on agentic traces is governed by whether the prefix cache holds rather than by peak token rate, and collapses past concurrency 8 on a B300 node. A single-stream tokens-per-second record and a multi-tenant agentic serving cost are close to unrelated numbers.

    The attached image is the Artificial Analysis "Output Speed" bar chart and it carries more information than the tweet does, so it is worth transcribing in full: Celeris-1 at 2038, Mercury 2 at 851, Gemini 3.5 Flash-Lite 365, Gemini 3.8 Flash 215, Command A+ 196, GPT-oss-120b (high) 193, Nemotron 3 Ultra 143, GPT-5.6 Terra (max) 126, Muse Spark 1.1 (xhigh) 124, Claude 4.5 Haiku 95, Inkling 86, Claude Sonnet 5 (max) 75, Gemma 4 31B 68, Mistral Medium 3.5 66, MiniMax-M3 65, Claude Fable 5 (with fallback) 62, GPT-5.6 Sol (max) 61, Grok 4.5 (high) 56, Qwen3.6 27B 56, Qwen3.8 27B 54, Claude Opus 5 (max) 54, Claude Opus 4.8 (max) 53, and Kimi K3 (max) last at 35. Reasoning models are flagged with a lightbulb icon and most of the chart carries one. Two corrections to the tweet's own framing follow from the chart. Celeris's "2x the next fastest" understates it, since 2038 against 851 is 2.4x. And the more useful reading is the shape rather than the leader: every frontier reasoning model sits between 35 and 75 tokens per second while the speed leaders sit 10 to 30x above them, which is the visual form of the claim that output speed and capability have become close to independent axes. Kimi K3 being dead last is the pointed detail given today's coverage of its architecture, and it is a defensible design choice rather than a failure: a model serving a 320:1 prefill-to-decode workload does not need fast decode.

  • AWS says its capacity is sold out into 2028 (@mattsgarman, Bloomberg). Garman, on Bloomberg Tech with Ed Ludlow, says "much of our capacity is already spoken for through 2027 and into 2028, and demand still significantly outstrips supply," framed as a reason AWS keeps investing rather than as a constraint. Two years of forward booking from the largest cloud provider is the clearest available statement that the scarce good in AI right now is not model quality but capacity, and it lands in the same digest as evidence that memory rather than compute is the binding constraint inside that capacity: the SemiAnalysis piece shows a B300 node holding Kimi K3 has only 3.25M tokens of KV budget after weights, and that prefix cache hit rate falls below 10% against a 95% theoretical rate the moment concurrency passes 8. A provider selling out two years forward and a serving stack that thrashes above concurrency 8 are the same shortage described from opposite ends.

  • Cursor hands coding agents write access to Google Workspace (@cursor_ai, changelog). New plugins let agents search and read mail, draft and send messages, apply labels and manage threads in Gmail; search, open, create and organize files in Drive; read schedules and create or update events in Calendar; and read, write and edit Docs and Sheets. Installed from the Customize page or the Cursor Marketplace. The reason this is worth more than a changelog line is the timing. Today's digest carries two papers establishing that agents fail specifically at noticing when shared state has moved: SWE-Touch measures a 7.7-point resolve-rate drop when a human edits task-relevant code mid-task, with agents retaining conflicting code or overwriting it without re-inspecting the repository, and ScrambleToolBench finds belief inertia when the environment drifts underneath a correctly-built model. A calendar, an inbox and a shared spreadsheet are mutable state that other humans change continuously and without notifying the agent, which is a harder version of the setting where the measured failure already appears. The write permissions are the part to note: read-only Workspace access would be a context-gathering feature, and write access makes every one of those five surfaces a place where a stale read becomes an action.

  • HUMAIN makes its first investment in a Saudi company (@TareqAmin_, announcement). Tareq Amin frames it with a line worth keeping: "Enterprise AI isn't defined by what you can prototype. It's defined by what you can deploy." The deal takes a stake in MOZN and pairs HUMAIN's full-stack AI infrastructure with MOZN's enterprise products and forward-deployed engineering, targeting sovereign AI for financial institutions and the public sector. No dollar figure is disclosed. The structural detail is "forward-deployed engineering," which is the Palantir model of putting your own engineers inside the customer, and it is the same shape as OpenAI's Presence offering logged in the 08-03 digest, where the stated fallback is that OpenAI engineers step in on complex cases. Two sovereign-adjacent enterprise AI plays in two days both pricing human deployment engineering into the product is a signal about how far from autonomous these systems actually are in production, and it sits oddly against the same week's record capital inflows.

  • Record capital into tech equity funds, with no thesis attached (@MarioNawfal). $75 billion flowed into global tech equity funds over five weeks, described as the largest five-week haul ever recorded, with $15.7 billion last week alone (third-biggest single week in history), a fourth consecutive week above $10 billion, and a four-week average of $14 billion that is 115% above the previous peak set in 2025 and 180% above the 2021 bull-market high. Attributed to AI enthusiasm, earnings hopes and rate expectations. Logged because the magnitude is the content and because it is the market-side companion to two research findings in today's digest that point the other way: yesterday's KV eviction result that on natural text the choice of eviction policy is nearly free because every policy collapses onto accumulated attention, and today's finding that linear attention's advertised constant-size cache is not constant in production. Algorithmic headroom is closing in exactly the areas where the capital is buying capacity, which is less a contradiction than a division of labour.

  • Three Robert Scoble meetings, one of which matters here (@Scobleizer on Goodfire, @Scobleizer on mixedbread, @Scobleizer on the Holodeck). The Goodfire lunch with founder Eric Ho and Curt Tigges is an interpretability company, which Scoble describes as tearing apart the LLM to understand how it thinks in order to build trust in AI systems. That connects directly to the wiki's standing finding, most recently in the 08-03 filler-token result, that hidden-state readouts are the only safety monitoring that has survived measurement while chain-of-thought monitoring has a demonstrated hole, so a company commercializing residual-stream interpretability is on the right side of that evidence. No technical claim is in the post itself; Scoble's substance is delegated to a Grok summary of his own videos, which is not a source. The mixedbread conversation is about a "new smart database for AI agents," with no detail beyond the category, and the category is real: UEmbed in today's digest is an attempt at the same problem from the model side rather than the storage side. The third is a 5,000 square foot Holodeck at Four Seasons Lake Austin that Scoble says John Carmack and Meta's CTO both flew to see, which is a genuine curiosity signal from two credible people and carries zero technical information in the post.

  • Kilo gives away Tencent's Hy3 for a week (@kilocode). Tencent's Hy3 is free inside Kilo Code for one week via VS Code or the Kilo CLI. Small, and logged for continuity, because Kilo is the vendor that has supplied this wiki with its two most useful production routing datasets (the plan/implement split lineage and today's 10,643-review code-review study), and a free-tier promotion of a Chinese open-weight model inside a routing product is the mechanism by which their open-weight share statistics keep moving. Their own site footer now reads "Kilo Joins Anaconda."

  • NVIDIA ships an agentic-commerce blueprint (@nvidia, NVIDIA). An open-source blueprint combining the ACP and UCP agent-commerce protocols in one codebase with the NeMo Agent Toolkit, Nemotron models and Milvus vector search, shipped ChatGPT- and Gemini-ready with GitHub and Docker Compose. The framing NVIDIA chooses is that merchants keep control over pricing, payments and compliance while AI becomes a primary interface for product discovery, which is a direct answer to the obvious retailer objection to agentic shopping. Two competing protocols supported from one codebase is the interesting engineering choice and the clearest sign that neither has won.

  • A Switch Transformer walkthrough, from the overnight slot (@ProfTomYeh, 08-03 evening). A 13-step hand-worked explanation of the Switch Transformer (Fedus, Zoph and Shazeer, 2022), the paper that made sparse mixture-of-experts practical at scale by routing each token to a single best expert rather than several. Included because it is the only piece of technical education in either overnight slot and because it is the direct pedagogical background to today's SemiAnalysis analysis of Kimi K3's LatentMoE, where the design decisions Yeh explains at the conceptual level (how many experts to activate, how large to make them) turn out to be set by an interconnect bandwidth budget rather than by anything about representation. Yeh's own framing is that MoE is how GPT-4, Claude, DeepSeek-V3 and Kimi all pack enormous parameter counts while activating only a slice per token.

  • Google DeepMind's post-AGI research team, and a question with no answer attached (@GoogleResearch, 08-03 evening). A booth Q&A at Deep Learning Indaba on the sociotechnical implications of AGI, sharing recent work from DeepMind's Post-AGI Research Team. Logged only because the existence of a team with that name at a frontier lab is itself the information; no claim, paper or finding is in the post.

  • Off-topic and personal, grouped. @mlevchin asking whether "load-bearing" is the next telltale em dash, which is a joke about LLM prose tics and the closest thing to AI commentary in his feed. @Tesla on FSD Supervised looking everywhere at once, pure marketing. @hexiang and @imagine both resharing the same Grok Imagine 1.5 reference tutorial from @alexutopia, still with no generated output attached, which is now the fourth consecutive slot where that feature gets enthusiasm and nobody posts a result. @Scobleizer noting a park is where X started 20 years ago. @ns123abc with four posts asserting the singularity has begun, containing no mechanism, timeframe or falsifiable content.

  • Skip. Roughly four fifths of the slot is outside scope. @MarioNawfal posts 20 items on Iran and the Strait of Hormuz, US polling, Panama Canal port operators, Telegram's brief worldwide App Store disappearance and return with no explanation from Apple, congressional insider-trading allegations, migration, and celebrity jiu-jitsu. @brivael posts 20 items of French and European political commentary plus Bill Gates and Fauci material, two of which attach a Senate PDF that the farmer captured as raw binary. @DoWCTO and @SeanParnellASW post US defence messaging. @MillionInt reposts a Napoleon's-march infographic. The @bayesiansapien curated repost feed was empty, which is now five consecutive slots.