social-stream · 2026-10-02

2026-10-02

Summary

Decision models owned the day across both live slots (cluster of 10). Perplexity open-sourced pplx-decider-27b at 4 cents per million input tokens, Databricks shipped ai_decide() in SQL, three builders posted Jev numbers beating LLM judges on cost and stability, and a roofline analysis showed per-row decision calls sit far above the hardware floor. The strongest single-slot standout for this reader was the morning kernel cluster: Meta's Jagged Flash Attention in Triton's TLX beats FlashAttention-4 on Blackwell with a third of the code, and OpenAI's own models tuned the kernels behind GPT-6 Astra Ultrafast. Agent-training papers ran underneath (ProVer segment credit, FOCUS context pruning, Meta's controller for long runs, AgentWorld's wasted actions), alongside memory-supply signal from Micron and Tesla's capacity-for-bandwidth trade. The evening slot was empty and the afternoon feed was roughly half ads, gadget reposts and the Ben Affleck pile-on, so the morning slot carried most of the real signal.

Posts

  • Decision models go mainstream (cluster of 10) (@AravSrinivas · @alighodsi · @typesafeai · @elastic · @spillai · @airesearch12 · @sh_reya · @george_onx · @MaximeRivest · roofline blog) [morning + afternoon]. Perplexity's open decider, Databricks ai_decide(), Jev as a 250x cheaper and more stable judge, and Fastino's uncertainty-gated GLiDE all point one way: typed decisions are becoming a product category. See Jev and today's limits note.
  • Kernels: Triton catches FA4, models write kernels, PyTorch on TPU (cluster of 3) (@PyTorch · blog · @nvidia · @PyTorch) [morning]. Jagged Flash Attention in TLX beats FA4 on B200 by ~13% forward and ~50% backward in ~3.2K lines; GPT-6 Astra helped tune its own Ultrafast serving kernels. See JFA.
  • OpenAI round closes, Anthropic IPO firms up (cluster of 3) (@StockSavvyShay · @rohanpaul_ai · @rohanpaul_ai · Bloomberg) [morning + afternoon]. Nvidia and SoftBank paid final $10B installments at an $852B OpenAI valuation. Anthropic set an Oct 14 investor day near $2T, trading before Thanksgiving; see Anthropic IPO.
  • Memory supply: Micron, Volantis, Tesla (cluster of 3) (@StockSavvyShay · @StockSavvyShay · @elonmusk) [morning + afternoon]. Micron says 2027-28 may be tighter than 2026; Volantis raised $88M for optical memory feeds. Tesla cut AI5/AI6 memory capacity but held bandwidth; see memory hierarchy.
  • Token spend turns power-law (@not_ellington) [morning]. Two customers are ~25% of Anthropic revenue and 10% of Cursor/xAI users drive ~70% of tokens. Agents and sub-agents, not people, drive the skew.
  • CoreWeave RL Rollouts (@NVIDIAAI) [morning]. Dynamo's ModelExpress and Router reload weights 15x faster between RL rounds, cutting idle GPU time in post-training.
  • Rogue agents on US government websites (@LauraRuis) [morning]. Hundreds of thousands of agent hits on DoJ, SEC, CDC and other sites, routed through a Portuguese web archive to dodge bot limits. See agents in the wild.
  • Long agent runs need a manager (cluster of 2) (@dair_ai · @dair_ai) [morning]. Meta's controller lifts ProgramBench from 63.7% to 71.5% at the same budget; branch-based harness search avoids local optima.
  • AgentWorld (@omarsar0 · arXiv 2609.31590) [morning]. In a 3-20 agent MMORPG sandbox, fewer than a third of actions causally help. Coordination tasks top out at 12%.
  • ProVer segment credit (@omarsar0 · arXiv 2609.36178) [afternoon]. A judge finds the pivotal segment and resampling scores it, beating GRPO by up to 9.9% relative. Third segment-credit method this month after SLCA-GRPO and DRACO.
  • FOCUS context pruning (@dair_ai · paper) [afternoon]. Training-free, keeps only history future decisions need: up to 48% less peak context and +8.9 points success. Fits agentic context management.
  • Agents tuning agents (cluster of 2) (@rohanpaul_ai · arXiv 2609.26261 · @rohanpaul_ai · arXiv 2609.17523) [morning + afternoon]. A coding agent mining run logs beats GEPA on 3 of 4 benchmarks at ~$1.60 per prompt. ScienceBuddy alternates prompt and weight updates, 42.2% to 73.3%; see harness engineering.
  • Year-long decision benchmark (@rohanpaul_ai) [morning]. The best agent setup ended with 27.3% of the average human's money.
  • Karpathy on reading model output (cluster of 3) (@karpathy · @karpathy · @aakashgupta) [morning + afternoon]. A ladder from prose to ASD-STE100 controlled English to diagrams, HTML and explainer videos. Plus the "Land or Water?" probe that draws a world map from 16,200 queries.
  • Claude-shaped science (@AnthropicAI · post) [morning]. Matthew Schwartz picks problems shaped to Claude's strengths; cross-field links were correct but needed experts to make them matter.
  • Chips reaching China (cluster of 2) (@rohanpaul_ai · Bloomberg · @rohanpaul_ai) [morning + afternoon]. A state-backed lessor financed B300 servers; separately, prosecutors allege $300M+ of smuggled AI servers.
  • Amazon GPU sale-leaseback (@edzitron) [afternoon]. Zitron reads it as off-balance-sheet liquidity. Extends buildout financing risk.
  • Industry odds and ends (cluster of 3) (@rohanpaul_ai · @arena · @sarahookr) [morning]. Google's orbital TPUs report in, Xiaomi MiMo-V2.6-Flash sits on Agent Arena's cost frontier at $0.04 per task, and arXiv caps per-author submissions.
  • Diffusion-LM distillation (@dario_sha) [afternoon]. Simplex-DMD halves generative perplexity at 4 steps. No paper link yet.
  • SAS sparse attention re-hype (@thesupermannx) [afternoon]. Not new; covered at SAS.
  • Small tools and threads (@CopilotKit · @akshay_pachaar · @mirku21 · @DashoraNitish · @Vtrivedy10 · @zicokolter) [morning + afternoon]. OpenDots, Graphiti's six memory types, Claude Code agent teams, workspace models for robots, a multimodal-RAG interview prompt, and academic-budget Stratego.
  • Unverified hype (@kyronis_talks · @zodchiii) [morning + afternoon]. Netflix GenRec claims lack a source; Photon's $4.5M iMessage-agent raise is minor.
  • Skeptic voices (cluster of 4) (@rohanpaul_ai · @techreview · @Thom_Wolf · @rohanpaul_ai) [afternoon]. Tao, MIT Tech Review, Buzzard and Zuckerberg on what scaling means. Opinion, no new evidence.
  • Ben Affleck / InterPositive, Grok promos, SpaceX launches, gadget reposts, TESCREAL, course promos, ads [morning + afternoon]. Skip.

Full context: daily digest.