social-stream · 2026-10-03

2026-10-03-morning

Summary

The morning slot's own scrape was empty again: no curated reposts, no tracked-handle tweets and no new bookmarks. The morning Following-feed capture carries the signal, and it covers the US evening of 10-02. The strongest signal is serving and kernels (cluster of 3): PyTorch and Red Hat put an autotuned Helion GEMM into vLLM that beats CUTLASS and DeepGEMM on Hopper, AMD previewed FlyDSL as an MLIR-native GEMM backend for TorchInductor, and Prime Intellect launched Prime Inference, a serving business grown from its own RL rollout fleet. The second cluster is the economics of the buildout (cluster of 3): Ed Zitron's long essay arguing AI's GDP contribution is almost all GPU sales and construction, a skeptical take on "annualized revenue" as a metric, and Morgan Stanley's view that power limits make Nvidia's compute per gigawatt the deciding number. The third is cost per task (cluster of 3): Agent Arena placed GPT-6.1 Sol at #5 for $0.56 per task and Sonnet 5.5 at #3 for $2.74, and OpenAI DevDay notes list prompt caching, reasoning effort, programmatic tool calls and batching as the main cost levers. Standouts outside clusters: Supabase's $150M round at $10.65B (70% of new databases are created by agents), Geoffrey Hinton's co-authored report on intelligence explosion, Meta opening Muse to hardware, and François Chollet's list of the three biggest shifts in AI (inductive over transductive, symbolic tool use, training on code execution). Late-evening reposts of KV-streams, ProVer, SkillAdam and the NVIDIA long-horizon paper kept circulating; they are covered in the digest.

Posts

  • Serving and kernels: Helion in vLLM, AMD FlyDSL, Prime Inference (cluster of 3). (1) PyTorch, Red Hat and Meta wrote one GEMM (matrix-multiply) kernel in Helion, PyTorch's kernel DSL, that can act as standard GEMM, Split-K (split the shared dimension across thread blocks for small outputs) or Swap-AB (swap operands for skinny decode-time matrices). An ahead-of-time autotuner picks the variant and settings per input shape. Inside vLLM's quantized linear backend on Hopper, it beats the default CUTLASS and DeepGEMM kernels across the tested models, with more than 10% throughput on some workloads (@PyTorch, blog). (2) AMD will present FlyDSL, a Python-native, MLIR-based kernel DSL plugged into TorchInductor's autotuning, with Triton-vs-FlyDSL results on Instinct GPUs at PyTorch Conference (@PyTorch). (3) Prime Intellect launched Prime Inference, serverless and reserved serving of open models on Blackwell. It began as the platform for its own RL rollouts, synthetic data and coding agents, processing nearly a trillion tokens a day internally, and runs NVIDIA Dynamo, vLLM, Mooncake (a shared KV-cache store) and FlashInfer. Its GLM-5.3 endpoint on OpenRouter reports near-zero tool-call errors since 09-22. The attached image is a launch title card ("PRIME INFERENCE" over a green network map); a stream of team reposts followed (@vincentweisser, blog). See Helion in vLLM.

  • The economics of the buildout (cluster of 3). (1) Ed Zitron's premium essay takes the WSJ and Brookings figure of AI at 3.63% of GDP over six years and argues six of the eight years are projections. Goldman puts AI investment at 1.9% of US GDP in 2026, but almost all of it is data-center construction and GPU sales, not companies renting compute or selling AI software. The ICT industry's share of nominal GDP has been flat for two years. His conclusion: when construction slows, AI services would have to cover the gap, and the data does not show that yet (@edzitron, essay). (2) A practitioner calls ARR (annualized revenue) "opaque and misleading" for inference businesses whose customers and GPU reservations can change within weeks (@not_ellington). (3) Morgan Stanley reinstated Nvidia as its top semiconductor pick, expecting agentic AI to drive more GPU demand and Rubin to reach an $80B run rate, because power constraints make compute per gigawatt the binding metric (@StockSavvyShay).

  • Cost per task on the agent leaderboard (cluster of 3). (1) GPT-6.1 Sol (Max) entered Agent Arena at #5 at a $0.56 median cost per task, within 2 points of GPT-6 Sol and Astra, 39% cheaper than Sol and 81% cheaper than Astra (@arena). (2) Claude Sonnet 5.5 (Max) debuted at #3 and is #1 in the Chat category, but at $2.74 per task it costs more than Opus 5.5 ($1.58) at a lower score, so it sits off the cost frontier (@arena). (3) Elvis Saravia's DevDay notes list the four cost levers OpenAI pushed: prompt caching, reasoning effort, programmatic tool calling and batch requests (@omarsar0).

  • Supabase raises $150M at $10.65B (@ycombinator). More than 13 million developers use it, it adds over 4 million databases a month, and 70% of new databases are created by agents or AI tools. A clean data point that agents are now the main customer of some developer infrastructure.

  • Intelligence explosion report from Hinton and coauthors (@geoffreyhinton, report). Hinton writes that recursive self-improvement long seemed far off, but many leading researchers now think it may come soon. Christian Szegedy reposted it. The capture holds only the landing page, so the report's specific claims are not summarized here.

  • Chollet on the three biggest shifts (@fchollet). His list: moving from transductive to inductive models (by far the biggest), symbolic tool use with code execution and harnesses, and training models on code execution and harnesses so they encode symbolic behavior. It lines up with the day's multi-harness RL story: the harness is becoming part of training.

  • Meta opens Muse to hardware (@alexandr_wang). Open-source ESP32 firmware and a Linux SDK let anyone build gadgets for Meta's Muse agent, plus a Home Link device for TVs and speakers. Cheap hardware and open tooling to seed an ecosystem early.

  • Sam Altman on Cerebras and on dot (@sama, dot). He calls Cerebras "a close partner" with a deep engagement on speed, amid speculation about the relationship. Separately he says OpenAI's dot assistant gets noticeably better each day as it learns his workflow. A same-night repost of Cerebras's CEO explains the wafer-scale pitch as a decode-bandwidth argument (@rohanpaul_ai).

  • Small models, argued with a release (@0xCodez, model). webAI's TwIL-LM3-Pro is a 3.6B model reported on par with Qwen3-8B on formal logic, 95.4% on BBH logic, 2.09 GiB at 4-bit and runnable on CPU. The post frames it with Karpathy's line that most calls never needed a frontier model. The numbers are the vendor's.

  • Cohere's new retrieval metric (@cohere). Embed 5 is reported with RCP-nDCG@10, which grades every returned result on actual relevance instead of crediting only items in the benchmark's answer key; Cohere says human reviewers back it. Worth knowing before comparing Embed 5 numbers with older nDCG tables.

  • Harness skills and courses (cluster of 3). Addy Osmani's agent-skills repo packages a senior-engineer workflow as skills, each with steps, the excuses agents use to skip them, and the evidence required before finishing (@undefinedKi, repo). A free harness engineering course was reshared (@khushiirl). Stanford's CS224V "Agentic AI" course from Monica Lam's group was reposted by Stanford NLP (RT).

  • Agents fixing their own harness bugs (@rohanpaul_ai). A paper from US and Chinese labs turns real bugs in agent harnesses (tool calls, memory, prompts) into a growing set of runnable tests, and finds coding agents miss most of them but improve with lessons from past fixes.

  • Skip. Market tickers, Elon Musk's SpaceX and robotaxi posts, the Schmidhuber-versus-the-Pope repost (its image is a meme), a Minecraft ballpark repost, cohere meetup registration links, and prompt-pack engagement bait.