social-stream · 2026-10-06

2026-10-06-morning

Summary

The morning slot had no new reposts or bookmarks. The signal comes from the small Following-feed capture that wrapped the US Monday evening, 12 posts, none with image attachments. The strongest single item is a clear explainer of chunked prefill in vLLM, the scheduler trick that keeps one long prompt from freezing every other user's stream. It pairs well with today's KV-cache papers, which attack the memory side of the same prefill-heavy traffic. The second is Leviathan, an open-source indexer that claims to hand an agent 436 tokens instead of 107,000 when searching a million records. There is a loose cluster of three posts on measuring agents instead of trusting them: an evals primer (a grader applied to a trace), an argument that AI engineering needs data-science discipline, and an X article on personal agents acting on stale private copies of company state. Two decision-model and RL items stand out for routing work: a practitioner's notes on training choice-order invariance into JEV-style System One models, and Tengyu Ma's claim that GRPO-style RL cannot be optimal. Reka's Rho-1 omni-model got a launch post. The rest is a single-source "AI rediscovered Newton's laws" thread, a LeCun repost and a career story.

Posts

  • Chunked prefill, clearly explained (@akshay_pachaar). Every request has two phases. Prefill reads the whole prompt in parallel, is compute-heavy, and builds the KV cache (the stored attention keys and values for every token). Decode then generates one token at a time and is limited by memory bandwidth, since each step re-reads the weights and the cache. If request A is streaming and request B arrives with a 32K-token prompt, processing B's prefill in one pass freezes A's stream. Chunked prefill splits B's prompt into token ranges and lets active requests decode between chunks, trading a slightly longer time-to-first-token for B against smooth streaming for everyone else. Useful context for today's CacheBack and Extender papers, which shrink what prefill has to build and store.

  • Leviathan: search a million records with 436 tokens (@joshuagunnn, repo). A single static binary that turns JSONL, JSON, CSV/TSV, SQLite or any database export into a ranked full-text index an agent can query. The author claims that at 1M records it gives the agent 436 tokens of context instead of 107,000, and finds the answer 99% of the time. These are the author's own numbers. The idea is the same as Microsoft's CorpusMap from the US afternoon (see the harness and agentic-search summary): structure the data once so the agent reads little.

  • Measure the agent, do not trust it (cluster of 3) (@businessbarista; @sermakarevich; @ashwingop). The evals primer defines an eval as a grader applied to a trace (the full record of one agent run) and lists grader families, starting with deterministic code checks before LLM judges. The second post argues LLMs "almost never fail loudly," so teams need held-out test sets and a metric before claiming a prompt or skill change helped. It singles out shared prompt and skill libraries that ship with no evaluation. The X article, only its opening captured, argues that connecting work apps to a personal agent (ChatGPT Dots, Instinct, Muse) creates a private copy of company state, so an agent can follow its permissions and still act against the company's latest decision. All three echo the US afternoon's theme of grading end state rather than the agent's own report.

  • Choice-order invariance for decision models (@neural_avb). An X article of notes from two weeks of building datasets and training "System One" decision models (the JEV-style models that return probabilities over a fixed set of answers). The title points at a known weakness: a model's choice can change when the answer options are reordered. Only the opening was captured. Relevant to routing because these models are now being sold as router and gating layers (see LLM routing).

  • "GRPO-style RL can't be optimal" (@stanfordnlp resharing @tengyuma). Tengyu Ma argues that with billions going into self-play and recursive self-improvement, the underlying RL methods need to improve, and GRPO (group-relative policy optimization, which scores each sample against its group's average) is not the end point. Only the first line was captured.

  • Reka Rho-1 launch post (@MateuszOnAI). Pitched as one model that understands, generates, reasons and acts across images, video and text with "one state, one loop, no tool calls." Per The Decoder it is a 19B omni-model trained on 320 H100s in about three months. It is a single-network alternative to routing between specialist models.

  • AI rediscovers Newton's laws from coordinates (@AnatoliKopadze). A thread on a Peking University system given only space and time coordinates from 46 mechanics experiments. It reportedly invented mass, then force, recovered Newton's second law, energy conservation and gravitation, and checked that two ways of measuring mass agree. It deliberately avoided language models. Single-source thread with no paper link captured, so treat as unverified.

  • Skip (@ylecun repost of his ETH Zurich "don't work on LLMs in academia" talk, already circulated yesterday; @Unnati_builds24 career story; @abeirami one-line reaction).