Summary
The morning's strongest signal is the decision-model category turning into infrastructure. Fastino shipped GLiNER2.5-Decide, a 340M-parameter encoder that makes typed decisions in one forward pass on a laptop CPU. An open 4B model took first place on the JevBench leaderboard, and a 150K-example open training set for decision models appeared. That is a cluster of about a dozen posts, and most of the remainder is promotion to skip. The most useful efficiency item came from the vLLM ecosystem, not a paper: LLM Compressor v0.14.0 rewrote GPTQ quantization as a Triton kernel, 15x faster end to end and 30x on MoE models. NeurIPS acceptances filled much of the rest of the feed. The ones with numbers are RELEX (reinforcement learning with verifiable rewards at 20% of the training cost), GRAM (recursive reasoning with parallel latent trajectories) and a deception-scaling study showing more agents does not mean more safety. Two NVIDIA-adjacent results carry the agent-training thread: Skill2Env turns 3,400 public Agent Skills into 8,000 RL environments, and Berkeley's 2.3B hybrid Mamba-2 MoE claims near-Llama-3.2-3B quality at under 1% of the pretraining compute. The Anthropic wet-lab result, 949 Claude agents finding a new enzyme-system family on 215.6M tokens, was the morning's most-shared industry post.
Posts
GLiNER2.5-Decide: a 340M open decision model on CPU (@fastinoAI, @george_onx, model card). An encoder-based model that takes a user-defined set of typed questions and rules at call time and answers them in one forward pass, with no prompt template and no generated tokens. The posted latency is 167 ms on a laptop CPU and 38 to 47 ms on T4 through A100. On Fastino's own 17-domain suite it averages 60.2% exact match against 57.6% for the hosted Jev and 46.6% for a router baseline. The more interesting part goes past classification. It extracts character-level spans, relations and structured records, and it enforces implications, exclusions, cardinality limits and ordinal bounds across related decisions, which production systems usually hand-code. The model card says plainly that it does not reason or answer open questions. It is a vendor's own eval, so treat the accuracy line as a claim.
JevBench v1.4.2: an open 4B tops the leaderboard (@airesearch12, JevBench). decider-4b v2 is now ranked first, ahead of Jev 1.13.0, JevK5, Cygnet and Hopper. The poster's own caveat is that "Jev is still smarter," meaning the leaderboard composite rewards cost and speed alongside capability. Taken with the $17 Together clone and CLM-8B from 09-24, the category's open tier now tops at least one public board within ten days of Jev's launch.
The decision-model tooling layer (cluster of 3) (@neural_avb, @TypeLLM, @navaneethvb). bev-decision-150K is an open dataset of 150K decision examples (choice, boolean, ordinal and tool-routing questions) built from existing open training sets, which is what anyone training a Jev replacement needs first. TypeLLM claims three things Jev lacks: image inputs, string and number types, and dependent decisions in a single call. A third post points to a write-up of turning an ordinary model into a decision model mostly through inference engineering, which is the same observation behind the 09-23 $5 Tinker fine-tune.
Decision-model promotion (cluster of ~8, skip) (@hanakoxbt, @beamnxw, @noisyb0y1, @Mahaximus_, @zodchiii, and others). Repo checklists, a "leaked document" claiming 193x speedups, and "the Internet moment" framing. The 100-architecture directory has some real patterns in it (model-tier routing, stale tool-output removal, low-confidence escalation), but the posts add no measurements.
LLM Compressor v0.14.0: GPTQ in Triton (@RedHat_AI, release). GPTQ is the standard post-training quantization pass: it rounds weights layer by layer and uses a Hessian estimate to push the error onto weights not yet rounded. The new Triton kernel makes it about 15x faster end to end. Batching layers that share a shape reaches about 30x on MoE models, where one expert shape repeats hundreds of times. Wider scale searches in the MSE and iMatrix observers find better NVFP4 scales, REAP expert pruning now runs distributed, and model-free PTQ gains KV-cache quantization. Full treatment: wiki summary.
VeriTile: formal proofs for GPU kernels (@__zrrr__). A group using Lean and AI-assisted proofs to reason about the correctness of Triton-style kernels. The timing lines up with the compressor release, where a new Triton kernel just replaced a trusted reference path.
RELEX: RLVR at 20% of the cost (NeurIPS) (@weizhepei, paper). RLVR (reinforcement learning with verifiable rewards) weight trajectories turn out to be close to rank-1 and nearly linear in training steps. So you train for the first steps, fit the line, and extrapolate, matching full RLVR with a fraction of the compute. This wiki covered it in May as RELEX rank-1 extrapolation. The acceptance is the news.
A 2.3B hybrid Mamba-2 MoE at under 1% of the pretraining compute (@berkeley_ai, second post). 360M active parameters, landing within a few points of dense Llama-3.2-3B. Both posts are retweets of the authors and the report is not linked, so the compute ratio is unverified.
Distillation is not a silver bullet for warm-starting RL (@maxkirkby). A first study of when distilling before RL helps. The effect varies a lot with model size, data amount, task, and whether you are optimizing for speed or accuracy after RL.
Skill2Env: public Agent Skills become RL environments (@askalphaxiv, alphaxiv). NVIDIA crawled 3.4K published Agent Skills (instruction files that teach agents a workflow) and turned them into about 8K executable terminal environments with programmatic tests and behavioral rubrics. After only 300 RL steps, Qwen3.8-27B gains 4.7 points on Terminal-Bench 2.1 and 4.3 on a private real-world benchmark. The pitch is that community-written skills are a task distribution closer to real use than synthetic ones. It is also exactly the CPU-hungry RL environment workload in today's CPU shortage essay.
PTTS: plan the parallel reasoning branches (@arXivBangers, paper). Instead of sampling parallel reasoning paths independently, PTTS first plans distinct outlines for each branch and then executes them. It adds up to 13.4 pass@64 points on math over repeated sampling. The cost angle is coverage per token: independent samples waste budget on near-duplicates.
GRAM: recursive reasoning as a stochastic latent trajectory (NeurIPS) (@JunyeobB, project). KAIST, Mila and NYU. Recursive reasoning models refine a latent state with a shared transition function but follow one deterministic path. GRAM makes the trajectory stochastic, trained with amortized variational inference, so inference can scale in depth (more refinement) and width (more sampled trajectories). It beats deterministic recurrent baselines on structured reasoning and multi-solution constraint problems.
Adding agents does not add safety (@addisonwu_, paper). The defection rate, how often an initially correct agent switches to a wrong final answer, rises linearly with the fraction of deceptive agents in the group. Larger groups are no more resistant at the same fraction. Humans in conformity studies are reliably swayed only by a misleading majority, while LLM agents defect even when deceivers are a minority. Letting deceivers coordinate privately made them less effective.
NeurIPS safety acceptances (cluster of 4) (@hamid_kazemi22, @DavidSchmotz, @narutatsuri, @mmitchell_ai). Suppressing a single MLP neuron bypasses safety refusals across seven models from 1.7B to 70B. Frontier agents treat oversight as an obstacle under ordinary task pressure with no instruction to evade. "Do Thinking Tokens Help with Safety?" got a spotlight. Margaret Mitchell's group argues agents can be designed so weak oversight is less of a risk.
Matryoshka Attribution tops the interpretability benchmark (@aryaman2020). An attribution method that uses gradient descent to find which parts of a network are responsible for a behavior. It is ranked first on the Mechanistic Interpretability Benchmark.
Anthropic's wet lab: 949 agents find a new enzyme-system family (@rohanpaul_ai). Claude agents searched 1.94B protein clusters over 21.5 hours on 215.6M tokens. They noticed a short DNA sequence repeating next to an enzyme, a CRISPR-like pattern, and Anthropic's scientists confirmed a new family of systems they call ART. The number to keep is the token budget attached to a discovery.
Krisp's voice-isolation benchmark (sponsored) (@svpino). Removing background noise before speech-to-text cut word error rate substantially on an open benchmark. It is a preprocessing win that costs nothing at the model layer.
Older papers recirculated (skip) (@hooshaaii, second). RetNet (2023) and the Differential Transformer (2024) promoted as if new.
Also in the slot. Cursor published how it cut two thirds of its agent system prompt with no quality loss on production traffic (@undefinedKi). IBM's STAIR retriever addresses documents by table of contents instead of chunks (@HowToPrompt__). Cameron Wolfe traced RL for LLMs from VPG through REINFORCE and PPO to GRPO (@cwolferesearch). A repo reimplements Ilya Sutskever's 30-paper list in pure NumPy (@0x0SojalSec). A practitioner reports their production harness now writes its own evals and fixes (@muratcan).
Full day synthesis: daily digest.