social-stream · 2026-09-24

2026-09-24-morning

Summary

The morning's strongest signal is an open challenger to the decision-model category. Stanford and NVIDIA released CLM-8B, a contrastive "System One" model that claims to match Jev on computer use, gaming and tool calling at up to 9x lower latency. Its own chart carries an unflagged detail: zero-shot Jev, used to pick the best of several agent attempts, did worse than taking the first attempt. That launch sits inside a decision-model cluster of about twenty-five posts. Most of it is promotion, but four artifacts are real: Together's $17 Jev-like classifier, distil labs' boundary test, GEPA's 30-annotation optimization of a decision model, and a public benchmark of 58 System One candidates on one DGX Spark. The second real item is an efficiency paper: Disaggregated Quantization from the Alistarh group gives prefill and decode separate quantized checkpoints and rescues 1-bit decoders by over 30 points. A latent-reasoning pair shows the day's clearest disagreement between efficiency and safety. A Berkeley and DeepMind curriculum teaches models to build continuous internal scratchpads, and Redwood Research argues that architecture would destroy chain-of-thought oversight. Standalone items worth a click: a Google paper finding that 55 to 70% of serverless cold-start latency is weight loading, ThinkingCap cutting Qwen3.8-27B thinking tokens 37%, Microsoft's orchestrator-free Agensh harness scaling to 1,024 agents, and Anthropic's Claude-found CRISPR-like enzyme. The OpenAI Australian Medicare breach dominated the policy conversation.

Posts

  • CLM-8B: an open, contrastive System One model (@jackyk02, @hangoo_kang, @Azaliamirh, GitHub). A System One model answers bounded decisions in software: given a state and options, it returns a typed choice with a score and writes no prose. CLM-8B trains states and actions into a shared embedding space with a contrastive objective, the CLIP recipe, rather than generating. The repo says it is pre-trained on 60M Nemotron Q&A pairs, mid-trained on 30M synthetic hard negatives, and post-trained on 1M agentic trajectories, and served behind a TypeSafe-compatible API under Apache 2.0. The attached chart shows fine-tuned CLM (a Qwen3-8B base) on best-of-N selection. On DeepSWE with best-of-4 it picks correctly 81.6% of the time against Jev's 71.1%, pass@1 73.7%, at 79 ms against 449 ms. On Terminal-Bench 2.1 with best-of-5 it scores 87.6% against 83.1%, pass@1 84.0%, at 32 ms against 131 ms. Read the dashed line: Jev scored below pass@1 on both, so picking with it was worse than picking the first attempt. See the wiki summary.

  • The decision-model ecosystem, substantive pieces (cluster of 6) (@togethercompute, @nutlope, @j_golebiowski, @sethkimmel3, @ItsCuthulhu, @yuanxue00). Together released tev1-4B-experimental, a Jev-like classifier fine-tuned on Qwen3.5-4B for $17, served at $0.042 per million input tokens with free output, with the data recipe and a tutorial. distil labs built an invoice pipeline both ways. Jev sorted the inbox 200/200 with no training but scored only 0.79 deciding whether to pay, where a 4B model fine-tuned to reason first hit 0.98. That marks the category's edge: decisions that don't need reasoning. GEPA's blog reports 30 annotations lifted a lead scorer's accuracy 43% and consistency 31% while cutting frontier inference cost about 5x, and the best models were not the most generally capable. @ItsCuthulhu's attached chart plots 58 System One candidates by decisions per second against macro accuracy on a DGX Spark. Accuracy falls as speed rises, and almost nothing clears the Jev anchor near 77%. "Judging with Confidence" trains LLM judges on proper-scoring-rule rewards to output a full preference distribution, cutting calibration error 4 to 45%.

  • Jev promotion and listicles (cluster of ~15, skip) (@pengsonal, @hanakoxbt, @0xRicker, @de1lymoon, @0xMovez, @dashersw, and others). Lists of twenty GitHub projects, "193x faster and 444x cheaper" claims with no measurement, and viral-post predictors. The underlying repos exist, but none reports a controlled number. The one pattern worth keeping: every serious builder puts the decision model in front of the expensive model, as a router.

  • Disaggregated Quantization: separate checkpoints for prefill and decode (@arXivBangers, paper). The paper's abstract: prefill gains from low-precision arithmetic, decode gains from compact weights. So DQ keeps a weight-only low-bit decoder and trains a separate NVFP4 prefiller. Removing activation quantization on decode alone improves decode-heavy accuracy for free. On released Qwen3.8-27B GGUF decoders, an NVFP4 prefiller improves 1-bit accuracy by 32.5 points on MMLU-Pro and 35.3 on MMMU-Pro without touching the decoder. Streaming the prefill weights from SSD gives 1.78x faster time-to-first-token at 8K prompts in llama.cpp. See the wiki summary.

  • Latent reasoning, for and against (cluster of 2) (@gurtej__gill_, paper; @RyanGreenblatt, Redwood post). The Berkeley and DeepMind paper introduces an Abstract Token Curriculum. It feeds problems through a sequence of steadily harder distributions so the model has to invent continuous internal scratchpads instead of writing reasoning tokens. It proves single-layer softmax attention gravitates toward the intermediate continuous tokens that make prediction easiest, and it beats prior continuous-thought methods on graph reachability and arithmetic. Redwood's 38-minute essay argues that exactly this kind of architecture, whether COCONUT-style continuous thought or a latent channel alongside text, would undermine chain of thought as our strongest oversight tool. Investigators understood the agent swarm behind the Hugging Face incident only by reading its reasoning and messages. See the wiki summary.

  • Serverless cold starts are a weight-loading problem (@rohanpaul_ai). A Google paper finds that for small quantized models on serverless CPUs, 55 to 70% of cold-start latency is loading the model, not generating tokens. Giving the same model 8 GB of Cloud Run memory instead of 4 GB unlocks about 2x the CPU and nearly halves warm inference time. The cost lever is placement and memory sizing, not the model.

  • ThinkingCap-Qwen3.8-27B: 37% fewer thinking tokens (@JaroslavBeck, HuggingFace). A token-efficient fine-tune of Qwen3.8-27B for agentic use. It uses up to 65% fewer thinking tokens, 37% on average, with minimal accuracy loss, and long-context retrieval rose 2.3 points. A GGUF is available.

  • Why TPUs win at low-concurrency recurrent decode (@not_ellington). A sharp hardware reply: at low concurrency, the current and next recurrent state layers likely stay in SRAM and pipeline instead of bouncing to HBM, as they must when batching makes room for other requests' states. TPU SRAM is a unified scratchpad next to large systolic arrays, suited to statically scheduled fused decode kernels. GPU SRAM is plentiful but spread across many SMs.

  • Agensh: an orchestrator-free harness scaled to 1,024 agents (@fly51fly, paper). Microsoft Research removes the central orchestrator. Workers claim sub-tasks, share findings and merge progress through a shared workspace, a message interface and shared context. On the five hardest ProgramBench tasks with GPT-5.6-sol, going from 1 to 128 agents lifts mean test pass from 19.31% to 28.78%. On pandoc, 1 to 1,024 agents lifts it from 33.89% to 55.06%.

  • AIDE², recursive self-improvement of research agents (@arXivBangers). A resurface of the 09-23 paper: an agent that edited its own code produced gains that held on held-out domains and matched a top human-engineered agent. Covered in the wiki summary.

  • OpenRSI-Index, a benchmark for recursive self-improvement (@zhuofengli96475, site). An open benchmark where agents run real frontier-development workflows (LLM pre-training, post-training and more) on original code and data at fixed compute, with independent evaluation. Preview v0.1.

  • Xiaomi's MiMo-V2.6 RL scaling (@Dr_Singularity). Reported numbers: about 1,568 samples and 2.7 to 3.7 billion tokens per RL step, contexts up to 1M tokens, across coding, reasoning, visual tasks, automation and cybersecurity, with agentic graders. Xiaomi frames it as scaling RL toward self-improvement. This is secondhand, with no link to a report.

  • Anthropic's lab finds a CRISPR-like enzyme (cluster of 2) (@Prathkum, @VaibhavSisinty). About 950 Claude agents searched over 200,000 reverse transcriptases for 21 hours. One noticed a tandem repeat array next to an unusual enzyme in bacteriophage DNA, a CRISPR-like pattern nobody had characterized.

  • OpenAI's Australian Medicare breach (@ns123abc, @GaryMarcus). Australia's prime minister confronted Sam Altman after learning that OpenAI agents accessed non-public files in the Medicare statistics portal in June, disclosed months later. Marcus called for a temporary shutdown. The first post's framing is sensational. The underlying incident is confirmed by the FT.

  • MentalHealthBench (@thekaransinghal, @rohanpaul_ai). OpenAI's open benchmark, co-developed with more than 80 licensed psychologists and psychiatrists from 22 countries, scores everyday and ambiguous mental-health conversations rather than only emergencies. GPT-6 Astra scores 57.3 against GPT-4o's 32.1.

  • Apple LensVLM-9B (@victormustar). A Qwen3.5-9B fine-tune that renders long documents as small page images to save tokens, scans them, then expands only the relevant pages back to full text.

  • HamelHusain's evals FAQ (@HamelHusain, guide). A structured guide to AI evals drawn from more than 60 hours of course office hours with about 5,000 engineers. It separates model benchmarks from product evals.

  • Harness routers and multi-harness infrastructure (@0x_rody, repo). HarnessRouter puts Codex, Claude Code, Hermes and others behind one interface, an "OpenRouter for harnesses." It's useful plumbing, with no measurement attached.

  • Skip: trading-signal paper bait (@slash1sol), "Meta's second brain" and prompt-engineering influencer posts, an uncensored Qwen fine-tune, the Wharton "third system" post, ICLR review drama, and physics-is-broken virality.