social-stream · 2026-09-18

2026-09-18-morning

Summary

The morning belonged almost entirely to one launch, and it is a launch with a direct line to the routing and cost questions this wiki tracks. Jev, the decision-only model from TypeSafe AI that emits structured choices and probabilities but never writes a sentence, dominated the slot with a cluster of eleven posts, and the important shift from its launch two days ago is that the posts are no longer vendor claims. They are practitioners reporting what they measured: a computer-use loop running at roughly 90 milliseconds per decision with no screenshots and no pixels leaving the machine, a game-playing demo costing under a cent, a Claude Code plugin that replaces the compaction summary with per-tool-call scoring, and a careful enumeration of where a near-free decision primitive slots into an agent stack (model routing, subagent orchestration, safety review, cache admission). The counterweight arrived in the same slot and is worth as much as the enthusiasm: @theo argued at length that the compaction plugin misunderstands what compaction is for, since it filters line by line without the thread context a summary would use. Outside the Jev cluster, the strongest single item is Bonsai 2 27B, a ternary-weight compression of Qwen3.8 27B posted by two separate accounts from PrismML, claiming 98.2% of the full model's benchmark performance at 5.95 GB against roughly 54 GB, running at 55 tokens per second in a browser on an M5 Max. Anthropic's disclosure that Claude now leads 26% of its own AI R&D propagated through several accounts with the numbers intact, and a recursive-self-improvement cluster of four posts formed around Google DeepMind's Dream-RSI and a general argument about what makes a self-improvement loop genuinely recursive. Sakana AI announced a new research group premised on the Transformer not being the end state. The quiet area this morning is anything about training: the entire slot is about inference, serving and scaffolding.

Posts

  • Jev practitioner reports (cluster of 11) (@milindlabs, @atomic_chat_hq, @da_fant, @_MaxBlade, @omarsar0, @sydneyrunkle, @hwchase17, @TIMNIRMAL, @SUOHA_AI). Jev is a model from TypeSafe AI, founded by Diogo Almeida, that takes application state plus a typed question and returns a decision over a predefined option set with a confidence score. It does not generate text at all. The morning's posts are the first wave of people who actually wired it into something. @milindlabs built computer use without any vision model in the loop. A local CoreML model segments every button and UI element on screen, on-device OCR reads the labels, and that text is all Jev receives. It returns a probability across the elements and names the one to click, then detection re-runs and it decides again. Roughly 90 milliseconds per decision, and no pixels leave the machine. The privacy property is incidental to the architecture but real. @da_fant's enumeration is the most useful post in the cluster because it names where the saving actually comes from. The headline case is subagent orchestration: long-running agents parallelize work across subagents, and every user message, email or subagent reply can wake the expensive orchestrator. He prices waking a frontier model with 100k input tokens at about $1, and the proposal is that a cheap decision model classifies each event and decides whether it warrants waking the orchestrator at all. That is admission control, and it is a much better fit for a near-free classifier than anything involving generation. @atomic_chat_hq ran it as a real-time controller, recalculating a safe tile every 330 milliseconds while obstacles fell, surviving 25 of 26 runs for under a cent. @_MaxBlade ran 50 games at once for the same order of cost. These are toys, but they establish the latency envelope, and @TIMNIRMAL drew the right generalization from them: the interesting claim is not classification, it is that a fast decision layer sitting under a slow reasoning layer is the shape a robot needs, where the large model understands the goal and something much faster decides whether to stop or dodge when the environment changes. @sydneyrunkle and @hwchase17, both from the LangChain side, framed it as a harness component rather than a model, which is the framing most likely to stick: constrained output is useful precisely because a harness is full of small decisions that currently cost a full LLM call. @SUOHA_AI surfaced the plugin that caused the day's argument, tamaratran/fast-jev-compaction, which replaces Claude Code's compaction summary by scoring every tool call and result in one fast request, dropping or truncating stale ones and keeping everything else verbatim. The post's own framing, that this is how a coding agent bypasses the conventional KV cache, is overreach, but the mechanism is real. Related wiki reading: llm-routing.md and the Jev page from 09-16.

  • The counter-signal on compaction (@theo). Worth reading in full against the cluster above. His argument has two parts. First, compaction is not a filter: its job is to clean up history so the agent stays focused, and it should be used sparingly when context gets long rather than constantly to keep context small. Second, the plugin decides line by line, per tool call, so it never sees the thread context that tells you what matters. A summarizer reads the whole conversation before deciding what to keep; a per-call scorer cannot. This is the sharpest technical objection anyone has raised to the Jev wave and it is not a hype complaint, it is a claim about what information the decision needs. Unresolved either way this morning.

  • Bonsai 2 27B, ternary weights at 98.2% retention (@pashakho, @evaninwords, @HamidMaei). PrismML compressed Qwen3.8 27B to ternary weights, meaning every weight is constrained to one of three values (-1, 0, +1) rather than a 16-bit float. The claim is 98.2% of the FP16 parent's benchmark performance at 5.95 GB against roughly 54 GB, about a ninth the size, with vision, tool use, agentic capability and long context retained, running at 55 tokens per second with WebGPU on an M5 Max. Three separate accounts from the team posted it within minutes, so treat the numbers as vendor self-report. The detail that carries the most information is generational: against the first Bonsai 27B two months ago the size barely moved while retention rose from about 95% to 98.2%, with the base model swapped from Qwen3.6 to Qwen3.8. The compression ratio is saturating and the quality retention is still climbing, which says the remaining headroom is in the recipe rather than the bit budget. See quantization.md.

  • Anthropic publishes its own recursive-self-improvement metrics (@ChrisGPT, @Skoorbkaz). The numbers travelled through the feed intact, which is unusual. Claude "leads" 26% of Anthropic's AI R&D work, up from under 1% in February, with the share at or above "AI collaborates" above 90%. Roughly 30,000 agents are doing research and engineering work inside the company at any one time, and over a billion agent decisions were reviewed in August, of which 0.002% were blocked by the online monitor. The offline system flags about 100,000 transcripts a week and escalates roughly 50 to human review. Anthropic's own stated reason for publishing is that it shows how close the world is to recursive self-improvement. @Skoorbkaz adds a detail with real consequences that most coverage skipped: the agents have persistent identities that survive model upgrades, which is what makes them auditable at all at that volume.

  • Recursive self-improvement without touching the weights (cluster of 4) (@Skoorbkaz, @Saboo_Shubham_, @TheTuringPost, @vartekxx). Google DeepMind's Dream-RSI is the anchor. The agent preserves its execution history, turns it into a replay world it can re-simulate cheaply, tries alternative strategies inside that replay, rewrites its policy, then redeploys. Reported: up to 162x fewer agent calls, 50x lower discovery budgets, and up to 2.09x better GPU kernel performance, with the weights left completely untouched. The finding worth carrying is that accumulated raw history beat summarized guidance, which cuts against the instinct to compress an agent's past into lessons. @TheTuringPost supplied the useful taxonomy in the same slot: a loop is only genuinely recursive to the extent that more of the loop becomes editable, from improving a model, to improving how it searches, to deciding what experience it needs next. @vartekxx summarized a Google document laying out a five-stage agentic engineering pipeline (specification, harness, trajectory, verification, meta-debug) whose closing instruction is the same thesis in plainer words: stop asking which model to use and start asking what system to build around it. See self-evolving-agents.md and agent-harness-engineering.md.

  • Sakana AI opens a research group on the premise that the Transformer is not the end state (@SakanaAILabs, post in Japanese, announcement). The Frontier Intelligence Group's stated position is that intelligence is not solved, and it picks three specific gaps against biological intelligence to work on: data efficiency, energy efficiency, and generalization. The announcement threads their existing work under that banner, including Continuous Thought Machines, learning by predictive coding, sparse Transformers, AI Picbreeder and Smart Cellular Bricks. The energy-efficiency framing is the one that intersects this wiki's interests most directly, and it is notable that a lab is making it a research target rather than a serving optimization.

  • Looped flows: thinking longer by refining hidden state instead of writing more tokens (@rohanpaul_ai). A Carnegie Mellon and Oxford paper arguing that a model can spend more test-time compute by repeatedly improving its hidden state rather than generating a longer chain of thought, which matters because generating tokens is the expensive part. The obstacle they address is that recurrent models are hard to train over long loops, since the hidden state updates become unstable or stop helping. Their fix is to train each update on a small denoising task while constraining the hidden state to remain useful for the next update. If this holds, test-time compute stops being synonymous with token count, which would change the accounting on every efficiency comparison in the area. See looped-transformers.md and test-time-compute-allocation.md.

  • Model collapse, restated for a general audience (@alex_verem). A popularization of the Oxford and Cambridge result that a model trained only on its own output degrades irreversibly across generations, with the mechanism stated correctly: every model slightly overproduces common patterns and underproduces rare ones, so training the next model on that output shrinks the tail again, and nine generations in the model has lost the plot. Not new, but it is circulating this week alongside the fair-use litigation and the argument that AI-generated text is contaminating the open web, and that pairing is what makes it worth noting.

  • Optical generative models (@CrazyShyyt). A claim that a generative model can run by physically passing light through structured optical layers, with only a tiny digital encoder producing the initial random seed, so the generative steps consume essentially no compute beyond the illumination source. The framing ("going to terrify Nvidia") is promotional and the post carries no link to the paper. Interesting if real, unverifiable as posted. Flagging rather than endorsing.

  • Bend2 enthusiasm (@rchaves). A programming language promising C-like speed, Go-like compile times, Rust-like memory management, Haskell-like typechecking, Lean-like proving, and GPU compilation. Ranked high on engagement but it is an aspiration list with no benchmark attached. Skip.

  • A quantum-field-theory preprint (@dr_logvinovich). High engagement, all-caps framing, a self-published Zenodo preprint claiming to overturn eighty years of quantum field theory. Off-topic and the presentation style is its own warning. Skip.