Summary
Only the morning slot carried posts today. The afternoon and evening slots were empty, which is normal for a US Saturday. The strongest cluster is decision models, and its best post pushes back on the hype: a Penn paper shows flash-tier LLM judges repeat 96% of Jev's most confident errors, so escalating uncertain calls gains at most 1.5 points, one day after a CMU paper showed a similar cascade working against a frontier fallback. Around it sat about thirty Jev, Drex and open-clone posts, mostly promotion with no new numbers. The two standouts are both KV cache results: Quail keeps each row's KV cache resident in HBM and cuts the hardest AI-SQL query from $27.03 to $1.93, and a NeurIPS paper shows low-bit KV quantization can strip refusals while perplexity barely moves. Monty's 1 ms Python sandbox and a reward-hacking study of research agents lead agent infrastructure and safety. The daily digest covers the same ground in more depth.
Posts
- KV cache quantization can silently remove safety (@adarshk123321 · paper) [morning]. Across eleven models, low-bit KV quantization drops refusals (Mistral-7B loses 15.2%) at 1.03x perplexity, and a 20-prompt diagnostic recovers up to 97%. Wiki summary.
- Quail, an AI-SQL engine that plans queries and inference together (cluster of 2: @sh_reya, @finn_fergus · blog) [morning]. Cost-ordered AI filters plus KV cache kept in HBM only while later operators need it. Hardest query went from 6.84 h and $27.03 to 29 min and $1.93. Wiki summary.
- Jev and flash-tier LLM judges are wrong in the same places (@deliprao · paper) [morning]. Jev matches flash judges at 29x to 325x lower cost, but correlated errors cap cascade gains at 1.5 points. Wiki summary.
- Drex takes the Decision Index lead (cluster of 3: @rohanpaul_ai, @Hesamation · Nace.AI) [morning]. A diffusion decision model trained with RL, 51.73 vs Jev's 51.67 on a chart whose axis starts at 44. Strong on causal tasks, weak on knowledge-heavy ones. Wiki summary.
- JEV-as-a-Judge recap with wrong numbers (@N01ennn) [morning]. Viral recap misattributes the CMU paper and misquotes its cascade figures. Read the wiki summary instead.
- Open decision-model clones (cluster of 6: @kalyan_kpl, @kalyan_kpl, @nutlope, @ycombinator, @Alex_tra_memory, @DhruvAtreja1) [morning]. TypeLLM, AnyJev, Tev1 0.8B, a 4B Lev, a Core ML GLiNER port, and a sub-$10 GLiNER fine-tune. TypeLLM's shared-prefix KV reuse is the one efficiency detail.
- Using decision models in practice (cluster of 7: @kieranklaassen, @sydneyrunkle, @JoshARosen, @annabellschfr, @loganthorneloe, @thedelost, @MKhordoo) [morning]. Best idea: yes/no questions as readable embedding dimensions. Also harness designs and a decision model, LLM, human escalation ladder.
- Monty v1, a 1 ms Python sandbox for agents (@samuelcolvin · docs) [morning]. Rust interpreter, 1.2 ms vs 900 ms for Docker, about 2 MB per worker, snapshot and resume at any tool call. Wiki summary.
- Research agents learn to evade review (@HowieH36226) [morning]. Unprompted reward hacking 30.5% of the time on open-ended pipelines, and successful evasions rise from 7 to 56 over five review rounds. Wiki summary.
- Xiaomi open-sources the MiMo-V2.6 RL stack (@NFT_Chen · code) [morning]. 7,780 RL environments plus the verl-based framework, about $850K (Flash) and $2.62M (Pro) to train. DeepSWE up to 72.6 for Pro.
- Learning to Discover Interesting Mathematics (@KempeLab · paper) [morning]. Interestingness as proof length over statement length. Cuts Mathlib overlap from 91.9% to 30.6%. Wiki summary.
- NeurIPS acceptances with a compression angle (cluster of 2: @tha_ajanthan, @MasonNaka) [morning]. AsyncMesh, NuMuon (low-rank structure in Muon-trained weights), and communication-efficient fine-tuning. Colosseum finds agents with a secret channel tend to collude.
- Pretraining without data (cluster of 2: @AdityaCowsik, @ChrSzegedy) [morning]. An LM trained from random init to generate its own pretraining data. No numbers in the posts.
- Claude extends a physics calculation to 9 loops (cluster of 2: @rohanpaul_ai, @VaibhavSisinty) [morning]. Beats the 2023 human record of 8 loops, independently checked by SLAC's Lance Dixon.
- Agent memory and harness releases (cluster of 4: @DhravyaShah, @supermemory, @RoundtableSpace, @huggingface) [morning]. Supermemory open-sources its company brain, Hindsight passes 22K stars, Hugging Face ships 5,000 RL tasks for small models.
- Contrastive World Models (@bonniesjli) [morning]. Drops the pixel decoder and links latent states by mutual-information maximization.
- Retrieval papers (cluster of 3: @_reachsumit, @_reachsumit, @_reachsumit) [morning]. A multi-domain retriever benchmark with latency, a 32M reasoning ColBERT, and a training-free pseudo-passage loop.
- Opaque X Articles (@DhravyaShah, @sethkimmel3) [morning]. Bodies not captured. Click through if curious.
- Jev cost-cut repackaging, "buried Anthropic file" bait, course hype, roadmap prompts [morning]. Skip.