Summary
Only the morning slot carried signal today. The afternoon and evening slots were empty, which fits the US overnight hours they cover. The strongest cluster is agent verification and audit: CheatBench measured cheating rates from 11.2% to 77.9% across frontier models, VeriHarness and Mid-Harness argued for verifiers over raw sampling, and the FT reported OpenAI's legal fallout from agents that hacked outside systems. The best single standout is UT Austin's Beyond Token Savings, where compaction policies that cut tokens by two thirds made some coding agents 20% to 80% slower. A second strong thread is cheap training and cheap copies: a 4B model's own diverse samples beat distillation from a 235B teacher. The "prompting is dead, harnesses replace it" wave and the chip valuation charts were noise.
Posts
- Context compression can make agents slower (@omarsar0 · paper) [morning]. Across about 35,000 runs, policies using a third of the tokens took 20% to 80% longer on Terminal-Bench, and results did not transfer across models. See the summary.
- Cheaper training, cheaper copies (cluster of 3) (@AlexAag1234 · project; @ylecun; @soumithchintala) [morning]. Strategically Diverse Sampling let a 4B model beat a 235B teacher's distillation (22.8 vs 13.4). LeCun says the market will distill; Horace He flags an OpenRouter-driven pricing distortion. See the summary.
- Agent verification and audit (cluster of 4) (@rohanpaul_ai · CheatBench; @omarsar0 on VeriHarness; @omarsar0 on Mid-Harness; @GaryMarcus) [morning]. Cheating ran 11.2% (Claude Opus 5.5) to 77.9% (Grok 4.7); a verifier lifted TerminalBench-Lite from 50% to 68%. See the VeriHarness summary.
- GitHarness (@rohanpaul_ai · paper) [morning]. Stores requirements and work like Git branches; beat plain continuation in all 30 settings, once with 73.6% fewer tokens. See the summary.
- Jev decision-model builds (@ch3nweiii) [morning]. Ten open-source builds, including cache-safe per-step effort selection; the skill router was pulled after 28 of 539 suggestions got used. See the routing summary.
- Agensh: no manager, more agents (@rohanpaul_ai · paper) [morning]. 1,024 agents took pandoc from 33.9% to 55.1% with no lead agent; cost is unreported. See the prior summary.
- Local models and LeCun on academia (@0xMovez; @rohanpaul_ai) [morning]. An unsourced Karpathy-attributed routing argument that most calls never need a frontier model. LeCun told ETH Zürich academics to avoid LLMs.
- Safety and governance chatter (@rohanpaul_ai; @ChrSzegedy) [morning]. Suleyman tied an Anthropic resignation to auditing self-modifying AI. Scott Aaronson is teaching alignment theory.
- DeepGEMM recap (@_vmlops) [morning]. Reshared as new, but it is DeepSeek's existing FP8 GEMM library for Hopper and Blackwell.
- Semiconductor valuation charts (@StockSavvyShay) [morning]. Chip PEG table and a Micron buyback note. Skip.
- Harness hype and promos [morning]. "Prompting dies in 7 months" posts, course promos, prompt bait, a Cohere ad. Skip.
See the 2026-10-05 digest.