Summary
The day had one dominant story and it sits directly on this wiki's cost-and-routing line. Jev, the decision-only model from TypeSafe AI that returns a structured choice plus a confidence score and never writes a sentence, took eleven of the morning's posts, and the shift worth noting is that these are practitioners reporting measurements rather than the vendor repeating its launch claims: a computer-use loop at roughly 90 milliseconds per decision with no pixels leaving the machine, a real-time game controller running for under a cent, and a clear-eyed account of subagent orchestration as admission control, where a near-free classifier decides whether an event justifies waking a frontier model at roughly a dollar per wake. The same slot carried its own best counterweight, and it is the sharpest technical objection of the day: @theo's argument that the Jev compaction plugin misreads what compaction is for, since scoring tool calls line by line never sees the thread context a real summarizer would use. The strongest single-slot standout outside that cluster is Bonsai 2 27B, a ternary-weight compression of Qwen3.8 27B claiming 98.2% of the parent's benchmark performance at 5.95 GB against roughly 54 GB, where the generational detail matters more than the headline: size barely moved since the first Bonsai while retention climbed from about 95%, so the compression ratio is saturating and the remaining headroom is in the recipe. Signal was unusually clean, with the recursive-self-improvement cluster around Google DeepMind's Dream-RSI and Anthropic's own 26% figure both arriving with numbers intact; the noise was easy to isolate, amounting to one unverifiable optical-compute claim, one benchmark-free language launch, and one off-topic physics preprint. The conspicuous gap across the whole day is training: every substantive item is about inference, serving or scaffolding.
Posts
- Jev decision-only model, first practitioner wave (@milindlabs · @da_fant · @atomic_chat_hq · @_MaxBlade · @omarsar0 · @sydneyrunkle · @hwchase17 · @TIMNIRMAL · @SUOHA_AI · wiki) [morning] (cluster of 11). Jev takes application state plus a typed question and returns a decision over a fixed option set with a probability, generating no text at all. The useful framing came from the LangChain accounts and from @da_fant: it is a harness component for the small decisions that currently cost a full model call, and its best case is admission control on when to wake an expensive orchestrator.
- Computer use with no vision model in the loop (@milindlabs) [morning]. A local CoreML model segments on-screen elements, on-device OCR reads the labels, and Jev picks the element to click from text alone at roughly 90 milliseconds per decision. The privacy property, that no pixels leave the machine, falls out of the architecture rather than being designed in.
- Jev as a real-time controller (@atomic_chat_hq · @_MaxBlade) [morning]. Recalculating a safe tile every 330 milliseconds while obstacles fell, 25 of 26 runs survived for under a cent, and 50 parallel games cost the same order. Toys, but they establish the latency envelope that @TIMNIRMAL correctly generalized to a fast decision layer sitting under a slow reasoning layer.
- Counter-signal: the compaction plugin misreads compaction (@theo · fast-jev-compaction) [morning]. Two claims: compaction exists to clean up history when context gets long, not to run constantly as a filter, and a per-tool-call scorer never sees the thread context that tells it what matters. This is a claim about what information the decision needs, not a hype complaint, and it went unresolved.
- Bonsai 2 27B, ternary weights at 98.2% retention (@pashakho · @evaninwords · @HamidMaei · wiki) [morning]. PrismML constrained every weight of Qwen3.8 27B to one of three values, claiming 98.2% of FP16 benchmark performance at 5.95 GB against roughly 54 GB, at 55 tokens per second on WebGPU on an M5 Max. Three team accounts posted within minutes, so treat as vendor self-report; the generational trend, flat size and rising retention, is the real finding.
- Anthropic publishes its own recursive-self-improvement metrics (@ChrisGPT · @Skoorbkaz) [morning]. Claude now leads 26% of Anthropic's AI R&D, up from under 1% in February, with roughly 30,000 agents running at any one time and 0.002% of a billion August decisions blocked online. The detail most coverage skipped is that the agents hold persistent identities across model upgrades, which is what makes auditing at that volume possible at all.
- Dream-RSI and self-improvement without touching the weights (@Skoorbkaz · @Saboo_Shubham_ · @TheTuringPost · @vartekxx · wiki) [morning] (cluster of 4). Google DeepMind's agent replays its own execution history as a cheap simulation, tries alternative strategies inside it, and rewrites its policy, reporting up to 162x fewer agent calls and 2.09x better GPU kernel performance with weights untouched. The result that cuts against instinct is that accumulated raw history beat summarized guidance.
- Sakana AI opens the Frontier Intelligence Group (@SakanaAILabs · announcement) [morning]. The stated premise is that the Transformer is not the end state, with three named gaps against biological intelligence: data efficiency, energy efficiency, and generalization. Energy efficiency as a research target rather than a serving optimization is the part that intersects this wiki.
- Looped flows: more compute without more tokens (@rohanpaul_ai · wiki) [morning]. A Carnegie Mellon and Oxford paper spends test-time compute by repeatedly refining the hidden state instead of generating a longer chain of thought, trained by a small denoising task per update that keeps the state useful for the next one. If it holds, test-time compute stops being synonymous with token count and every efficiency comparison in the area needs re-accounting. See also test-time-compute-allocation.md.
- Model collapse, restated for a general audience (@alex_verem) [morning]. A correct popularization of the Oxford and Cambridge result that training on your own output shrinks the rare tail each generation until the model loses the plot. Not new, but it is circulating alongside the fair-use litigation and the web-contamination argument, and that pairing is why it registered.
- Optical generative models (@CrazyShyyt) [morning]. Claims a generative model can run by passing light through structured optical layers with only a tiny digital encoder producing the seed, so the generative steps cost almost nothing. No link to a paper and the framing is promotional, so flagging rather than endorsing.
- Bend2 language launch (@rchaves) [morning]. C-like speed, Go-like compile times, Rust-like memory management, Lean-like proving, GPU compilation, and no benchmark attached. Ranked high on engagement, which is the only thing it has. Skip.
- Quantum-field-theory preprint (@dr_logvinovich) [morning]. A self-published Zenodo preprint in all caps claiming to overturn eighty years of physics. Off-topic and the presentation is its own warning. Skip.