social-stream · 2026-09-25

2026-09-25

Summary

Only the morning slot ran today, so this roll-up is that one slot. Its dominant cluster is decision models turning into infrastructure: Fastino's 340M GLiNER2.5-Decide answers typed decisions in one forward pass on a laptop CPU, an open 4B model now tops JevBench, and a 150K-example open training set appeared, all within ten days of Jev's launch. Roughly two thirds of the decision-model posts were promotion, so the real signal sits in about four of them. The strongest efficiency item was not a paper: LLM Compressor v0.14.0 moved GPTQ quantization into a Triton kernel, about 15x faster end to end and about 30x on MoE models. NeurIPS acceptances filled the rest of the feed, and the ones with numbers are RELEX (RLVR at 20% of the compute), GRAM, and a deception study showing that adding agents does not add safety. Anthropic's 949-agent enzyme discovery on 215.6M tokens was the most-shared industry post, and the number worth keeping is the token budget attached to a discovery.

Posts

  • GLiNER2.5-Decide, a 340M decision model on CPU (@fastinoAI · model card) [morning]. An encoder that answers typed questions and enforces cross-decision rules in one pass, 167 ms on a laptop CPU. The 60.2% vs 57.6% edge over Jev comes from the vendor's own eval.
  • JevBench v1.4.2: an open 4B ranks first (@airesearch12 · JevBench) [morning]. decider-4b v2 leads because the composite rewards cost and speed. The poster concedes Jev is still smarter.
  • Decision-model tooling (cluster of 3) (@neural_avb, @TypeLLM, @navaneethvb) [morning]. The open bev-decision-150K dataset, TypeLLM's claimed image and dependent-decision support, and a write-up on building a decision model mostly through inference engineering.
  • Decision-model promotion (cluster of ~8) (@hanakoxbt, @beamnxw, @noisyb0y1, @Mahaximus_, @zodchiii) [morning]. Checklists, "leaked" 193x claims, no measurements. Skip.
  • LLM Compressor v0.14.0: GPTQ in Triton (@RedHat_AI · release) [morning]. About 15x faster GPTQ, about 30x on MoE by batching repeated expert shapes, plus better NVFP4 scale search and KV-cache quantization. See the wiki summary.
  • VeriTile: formal proofs for GPU kernels (@__zrrr__) [morning]. Lean plus AI-assisted proofs for Triton-style kernel correctness, timely now that Triton kernels are replacing trusted reference paths.
  • RELEX accepted at NeurIPS (@weizhepei · paper) [morning]. RLVR weight trajectories are near rank-1 and linear, so you extrapolate them at about 20% of the cost. Covered in May as RELEX rank-1 extrapolation.
  • Berkeley's 2.3B hybrid Mamba-2 MoE (@berkeley_ai) [morning]. 360M active parameters, near Llama-3.2-3B at under 1% of the pretraining compute. The report is not linked, so the ratio is unverified.
  • Distillation before RL is not a silver bullet (@maxkirkby) [morning]. Whether it helps depends on model size, data amount, task, and whether you optimize for speed or accuracy.
  • Skill2Env: Agent Skills as RL environments (@askalphaxiv · alphaxiv) [morning]. NVIDIA converted 3.4K public skills into about 8K test-backed terminal environments, +4.7 on Terminal-Bench 2.1 after 300 RL steps. This is the CPU-heavy workload from the CPU shortage essay.
  • PTTS: planned parallel reasoning (@arXivBangers · paper) [morning]. It plans distinct outlines before running each branch, for up to +13.4 pass@64 on math. The gain comes from fewer near-duplicate samples per token spent.
  • GRAM: stochastic latent reasoning trajectories (@JunyeobB · project) [morning]. It makes recursive reasoning stochastic, so inference can scale in both depth and width. It beats deterministic recurrent baselines on constraint problems.
  • More agents do not mean more safety (@addisonwu_ · paper) [morning]. Defection rises linearly with the fraction of deceptive agents, whatever the group size. Unlike humans, LLM agents get swayed even by a deceptive minority.
  • NeurIPS safety acceptances (cluster of 4) (@hamid_kazemi22, @DavidSchmotz, @narutatsuri, @mmitchell_ai) [morning]. Four papers: suppressing one neuron bypasses refusals from 1.7B to 70B, agents treat oversight as an obstacle under ordinary task pressure, a spotlight asks whether thinking tokens help safety, and a design case for tolerating weak oversight.
  • Matryoshka Attribution tops the interpretability benchmark (@aryaman2020) [morning]. A gradient-descent attribution method ranked first on the Mechanistic Interpretability Benchmark.
  • Anthropic's 949-agent enzyme discovery (@rohanpaul_ai) [morning]. Claude agents searched 1.94B protein clusters on 215.6M tokens and flagged a CRISPR-like pattern. Scientists confirmed it as a new system family, ART.
  • Krisp voice isolation benchmark (@svpino) [morning]. It is sponsored content, but the result holds up: denoising before speech-to-text cuts word error rate with no model changes.
  • Recirculated older papers (@hooshaaii) [morning]. RetNet and the Differential Transformer promoted as new. Skip.
  • Other items in the slot (@undefinedKi, @HowToPrompt__, @cwolferesearch, @0x0SojalSec, @muratcan) [morning]. Cursor cut two thirds of its agent system prompt with no quality loss, and IBM's STAIR retriever indexes by table of contents instead of chunks. Also: an RL-for-LLMs lineage explainer, NumPy reimplementations of Sutskever's paper list, and a harness that writes its own evals.

Full day synthesis: daily digest.