social-stream · 2026-10-02

2026-10-02-afternoon

Summary

This slot had no curated reposts, so everything here comes from the 134-post Following-feed capture. About half of that capture is ads, robot-gadget reposts and SpaceX launch retweets. The strongest signal is a decision-model cluster (cluster of 7). Perplexity open-sourced pplx-decider-27b at 4 cents per million input tokens, and Databricks shipped ai_decide(). Three builders posted Jev numbers against LLM judges, all pointing the same way: typed decisions are far cheaper and more stable than generated text. Next is a run of agent-training and context papers: ProVer for step-level credit assignment, Microsoft's FOCUS for training-free context pruning, and ScienceBuddy, which alternates prompt and weight updates. On hardware, Musk said Tesla cut AI5/AI6 memory capacity but kept bandwidth constant, and Bloomberg reported B300 servers reaching China through a state-backed lessor. Skip the Ben Affleck pile-on, the Grok promos and the ads.

Posts

  • Decision models go mainstream (cluster of 7). (1) Perplexity open-sourced pplx-decider-27b, a multimodal decision model, with a Decisions API at $0.04 per million input tokens and free output tokens (@AravSrinivas). (2) Databricks shipped ai_decide(), which runs decision models as a SQL function over governed tables (@alighodsi · blog). (3) A You.com risk-monitoring agent used Jev as its judge. Report quality was the same, but the LLM judge gave one threat three different scores (0.35, 0.68, 0.50) and missed 5 of 11 investigations. Jev missed none and was 250x cheaper (@typesafeai). (4) Elastic used Jev to rerank hybrid search results: nDCG@10 went from 0.935 to 0.957 (@elastic · blog). (5) VLM-based document classification used a fixed 372 tokens per page, against a 2,002 median for OCR-to-text, about 7x fewer (@spillai). (6) JevBench now ranks on capability, capping each entry's cost and latency at 2x Jev (@airesearch12). This is routing and judging becoming a product category. Background: Jev, a decision-only model, its correlated-error caveat, today's limits note.

  • ProVer: credit for the step that actually mattered (@omarsar0 · arXiv 2609.36178). GRPO (group-relative policy optimization) gives every token in a trajectory the same advantage. ProVer has a judge compare successful and failed rollouts to find the pivotal segment. It then resamples continuations from just before and just after that segment, and uses the change in success rate as that segment's credit. The gain over GRPO is +9.9% relative at 2B and +7.1% at 4B, and it holds with a smaller judge. It is the third segment-credit method in a month, after SLCA-GRPO and DRACO.

  • FOCUS: drop the history your next decision doesn't need (@dair_ai · paper). This is Microsoft's test-time context compression for agents. It keeps only the past interactions that future decisions depend on. It needs no training, so it works in front of closed APIs. It cut peak context by up to 48% and raised success by up to 8.9 points over the full history. Less context scoring higher fits the agentic context management thread.

  • ScienceBuddy: alternate prompt edits and weight updates (@rohanpaul_ai · arXiv 2609.17523). It turns user corrections into scored test tasks. It then alternates between rewriting prompts and skills and retraining the model. Together the two lifted a science agent from 42.2% to 73.3%. With a 4B model, prompt and skill edits alone went from 31.1% to 51.1%. This is harness and weights co-evolving; see agent harness engineering.

  • Tesla trades memory capacity for volume, keeps bandwidth (@elonmusk, follow-up · @DirtyTesLa). AI5 dropped to 72GB of LPDDR5, later nudged up to 96GB, and AI6 to 144GB of LPDDR6. Bandwidth was held constant, and Musk's argument is that bandwidth, not capacity, limits inference. That is a direct read on DRAM supply pressure. The viral "10B robots × 200GB" post is napkin math, but it points at the same constraint. See memory hierarchy.

  • B300 servers reach China through a state-backed lessor (@rohanpaul_ai · Bloomberg). A leasing firm controlled by local governments financed 32 Asus servers built on export-restricted B300s. The evidence surfaced in filings lodged with the PBoC credit registry. Nvidia says it will investigate.

  • Amazon sale-leaseback of its chips (@edzitron, follow-up). Zitron reads Amazon selling GPUs to investors and leasing them back as off-balance-sheet liquidity. He asks why Amazon would keep buying GPUs if it needs to do this. It is opinionated, but it extends AI buildout financing risk.

  • Anthropic IPO timeline firms up (@rohanpaul_ai · Bloomberg). An investor day is set for Oct 14 at a valuation near $2T. Marketing could start the week of Nov 9, with trading before Thanksgiving, which means an S-1 by late October. OpenAI ruled out a 2026 listing. Also: Nvidia and SoftBank paid their final $10B each, completing OpenAI's $60B in pledges (RT). See Anthropic IPO.

  • SAS sparse attention, re-hyped (@thesupermannx). This is a hype-voiced re-share of Tencent's frozen-backbone gated sparsification on SGLang block-sparse kernels. It is not new; covered at SAS.

  • Distribution-matching distillation for diffusion LMs (@dario_sha). It introduces two few-step distillation methods: Simplex-DMD for very low step budgets and Reinforce-DMD for larger ones. At 4 function evaluations, Simplex-DMD halves the best baseline's generative perplexity at matched entropy. No paper link is in the post, so click through to read.

  • Workspace models: a latent harness for robots (@DashoraNitish). A memory architecture that turns a slow reasoning agent into a low-latency robot policy. Harness ideas are moving into embodied control. Click through to read the thread.

  • Netflix GenRec (@kyronis_talks). Watch history is serialized as text and fed to an LLM ranker, which beats the tuned production recommender with about 40x fewer labels. The thread is hype-voiced and has no source link, so treat the numbers as unverified.

  • Claude Code agent teams (@mirku21 · docs). Opus as the lead, Sonnet teammates working in worktrees, and Fable as an adversarial reviewer. It is a model-tier routing pattern inside one tool. The "official tip" framing is the poster's, not Anthropic's.

  • Karpathy's "Land or Water?" eval (@karpathy). Query an LLM 16,200 times with lat/long coordinates and plot the answers: a world map comes out. A neat probe of what compression stores. Related: ASD-STE100 controlled English as an anti-slop prompt hack (@aakashgupta).

  • Multimodal RAG as an interview question (@Vtrivedy10). A good prompt that drills into per-modality ReadFile tools, tokenization, and harness design. Useful for studying.

  • Skeptic and big-picture voices (cluster of 4). Zuckerberg says multi-GW clusters brute-force AGI (@rohanpaul_ai). Terence Tao says AI solutions "don't feel like intelligence" (@rohanpaul_ai). MIT Tech Review argues "LLMs don't reason" (@techreview). Kevin Buzzard asks what math is for if AI keeps scaling (@Thom_Wolf). Opinion, no new evidence.

  • Superhuman Stratego on an academic budget (@zicokolter). A two-year solo-ish project. Worth a click if you track game-playing RL.

  • Ben Affleck / InterPositive (cluster of 6). Netflix's $587M cash buy of a 16-person film-AI startup, plus a pile of memes. One line of signal: domain craft knowledge commanded the premium (@rohanpaul_ai). Skip the rest.

  • Grok in XChat, Grok Bot, SpaceX launches, Ronald_vanLoon gadget reposts. Skip.

  • Promoted posts (token2049, AnswerThis, CodeRabbit, GitLab, Qodo, UiPath, trading and shopping ads). Skip.