Efficiency shorts, 2026-10-08
Source: HuggingFace Daily Papers 2026-10-07, Kurate cs.LG weekly, and the X feed. Short entries for papers that touch attention cost, KV memory, latency or token budgets but do not need their own page.
- UNREAL (arXiv 2610.08463, raw). One mechanism for both RAG retrieval and long-context reading. It encodes chunks and builds retrieval queries from a frozen LLM's own hidden states, adding under 500K trainable parameters. On a 21M-chunk Wikipedia index it beats retriever-reranker stacks (HotpotQA recall 49.1% to 73.2%). On long context it removes distractors before generation: NoLiMa at 128K goes from 1.0% to 24.8%, and FLOPs and time-to-first-token fall versus full-context reading from about 32K tokens up. Same thesis as Periscope (10-07): select evidence first, never pay full attention over everything.
- HLA, hybrid linear attention (arXiv 2610.05842, raw). Gated DeltaNet compresses history into a fixed state and loses sparse, distant facts. HLA stores each finished chunk as an exact affine state update and lets each query gate which chunk updates to apply. Qwen3.5 0.8B-9B: up to +5.6 on LongBench-V2, +4.0 on RULER; gains grow with length (0.8 at 4K to 4.2 at 32K, from scratch at 1.3B). The language-model sibling of yesterday's HLA-WM.
- Multilinguality in hybrid attention LLMs (arXiv 2609.35378, raw). Cross-lingual alignment spikes at the first full-attention layer. Every alternative layer ordering beat the standard one in multilingual distillation, learning up to 2.5x faster. Suggests hybrids should start with a full-attention layer.
- DeCoPrune (arXiv 2609.39096, raw). KV pruning for streaming video diffusion. Keeps tokens that were hard to denoise (they carry unpredictable visual evidence). Prunes 85%+ of historical KV, 4x faster continuation, near-full long-range recall on a new 58-episode benchmark.
- Speculative tool execution for voice agents (arXiv 2610.07641, raw). Predict tool calls from partial speech recognition, run them early, cache results. Median time-to-first-audio 5.79 s to 4.60 s on an Android assistant, with worst case bounded by the serial pipeline. Speculative decoding's idea applied to tools.
- Fixed input embeddings (arXiv 2610.04002, raw). 1.7B models trained on 100B tokens with fixed 16-bit token-ID codes instead of a learned embedding table still reach 52.4% HellaSwag. The learned table is better, but not necessary. Removes 100.7M parameters.
- Looped transformers, three machines (AlphaSignal, via X). A useful classifier: depth loops, time loops, flow loops. Nanbeige4.2-3B (Apache 2.0) runs 22 layers twice: about 2x block compute and 2x KV cache for a 4B file, keeping ~75% of standard token efficiency. An independent 140M test of the viral Recurrent Looped Transformer lost to a plain transformer (3.688 vs 3.642 nats) at ~20x the GPU-hours. See looped transformers.
- SPIN: Shadow Predictive Indexer for sparse attention (arXiv 2610.09025) and Jumping the Line: exploiting length predictions in LLM scheduling (arXiv 2610.03430) appear in this week's Kurate leaderboards (cs.LG #8 and cs.AI #9). This week's tournament has not scored yet, so their ranks carry no quality signal; noted for follow-up.
Related
KV cache · Attention mechanisms · Speculative decoding · Looped transformers