Efficiency shorts, 2026-10-07
Source: HuggingFace Daily Papers 2026-10-06 (plus one X-feed paper). Short entries for papers that touch memory, attention cost or token budgets but do not need their own page.
- HLA-WM (arXiv 2610.05739, raw). Gated DeltaNet (a linear-attention layer that compresses history into a fixed-size state) forgets distant scenes in long video world models. HLA-WM caches compact per-chunk transition summaries, retrieves the relevant chunks by camera geometry, and recomposes a query-specific recurrent state. Training-free. 12x less history memory than full KV caching at 60 s, at most 1.6% throughput loss, and better revisit consistency (+0.74 dB PSNR, -28.5% rotation error). A retrieval-addressable recurrent state is a middle path between a full KV cache and a lossy fixed state.
- Prism (arXiv 2610.05416, raw). Dynamic block-sparse attention for training joint video-audio models at 2K. Blocks are shaped per spatiotemporal zone from visual variance and audio-to-video attention strength, with per-query sparsity. 2.5x training speedup over full attention with better quality.
- CANOPY (arXiv 2610.00923, raw). Post-retrieval evidence compression for multimodal RAG: items become hierarchies, a fine-tuned node encoder scores regions, and parent-relative refinement keeps regions at mixed granularity with no LLM calls. 14.2-27.7% fewer reader tokens at equal accuracy; most accuracy gains come from critic-triggered extra retrieval.
- The cost of agentic retrieval (arXiv 2610.05750, raw). A ReAct loop over the same embedding model adds +8.7 nDCG@10, but costs 107.4 s vs 0.67 s and ~764K input tokens per query. The cleanest public price tag yet for "let the agent search" and a direct target for decision-model gating (see SearchJev).
- Foresight (arXiv 2610.03123). Streaming VLMs plan which future frames need computation, adapting compute to scene dynamics without retraining.
- Long-video joint allocation (arXiv 2610.04318). Lower per-frame resolution buys denser temporal coverage at a fixed visual-token budget; treat frames, pixels and frame count as one allocation problem.
- PAIR (context compression for agents, arXiv 2609.36526, via X). Replays an agent from the same state with and without one compression step, so randomness cancels. A few compressions cause the big failures by dropping an unresolved task condition; rewriting only the matching part of the compression prompt nearly matches no-compression accuracy on AppWorld, OfficeBench and tau-Bench. Extends Beyond Token Savings (10-05), which found compaction at a third of the tokens can run slower.
- The Numerical Linear Algebra of LLMs (arXiv 2610.04631). A survey for numerical-methods specialists: the matrix and tensor concepts under LLMs and NLA's recent contributions. A reading list, not a result.
- PaLoRA (arXiv 2610.04226). Replaces the fixed small learning rate in LoRA continual learning with a paced, theory-guided gradient scale.
Related
KV cache · Attention mechanisms · Test-time compute allocation