Compaction is a bet: REMORY, incremental deep research, and the swarm essay (2026-10-10)
Sources: HuggingFace Daily Papers 2026-10-09: REMORY, Learning Residual Memory for Context Compaction (arXiv 2610.11287, raw); Incremental Open-Ended Deep Research with Structured Harness (arXiv 2610.11566, raw). Prime Intellect, On the Nature of the Swarm (Konstantin Dunas, via @vincentweisser) and the Prime Agent Rust rewrite (@PrimeIntellect), X feed. Abstracts and article text.
TL;DR. Agents now run far past their 1M-token windows by compacting: summarize, hand off to a fresh context, continue. Prime Intellect's essay names the cost: every compaction is a bet on what will matter later, and the summarizer must decide before it knows. Writing notes to disk turns the question from "what must I remember" into "what must I know exists," but re-reading costs the very context that ran out, and whatever the previous agent understood is gone. The essay's conclusion: inference scaling ends in many agents (sub-agents and swarms), and Prime Agent demonstrated it by orchestrating 2,000+ agents across 10,000+ sandboxes for two weeks to rewrite itself in Rust. REMORY attacks the bet directly: a small network writes a bounded set of soft memory tokens appended after the text summary, trained so a frozen LLM's continuation matches what it would produce with full history. It approaches full-context quality on SummHay using 5.2% of input positions, and cuts repeated tool outputs and tool errors on BrowseComp and Terminal-Bench 2.1 for Qwen3.8-27B and GLM-5.3-Flash. Incremental-OEDR keeps a research report as a structured evolving state and updates it instead of regenerating: 33% fewer tokens and 61% fewer search calls on DeepResearch Bench.
flowchart LR
H["Full history<br/><small>past the window</small>"] --> S["Text summary<br/><small>lossy, chosen early</small>"]
H --> N["Memory network<br/><small>REMORY</small>"]
S --> N
N --> M["Soft tokens<br/><small>bounded, appended</small>"]
S --> F["Frozen LLM<br/><small>continues the task</small>"]
M --> F
F --> O["Next actions<br/><small>fewer repeats, errors</small>"]
classDef input fill:#d0ebff,stroke:#1971c2,color:#1b1b1b,stroke-width:2px
classDef core fill:#e5dbff,stroke:#6741d9,color:#1b1b1b,stroke-width:2px
classDef loop fill:#fff3bf,stroke:#f08c00,color:#1b1b1b,stroke-width:2px
classDef exit fill:#d3f9d8,stroke:#2f9e44,color:#1b1b1b,stroke-width:2px
class H input
class S,M loop
class N,F core
class O exit
How it relates to the wiki
- Answers the 10-04 open question from a different angle. AutoCompact (10-04) learned when to compact (+9.2 SWE-bench Verified). REMORY learns what survives beyond words. Just-in-Time Memory (09-24) argued to defer curation until the question is known; REMORY instead makes the early summary less lossy. Both reject blind compression. See Agent memory.
- Cost link to KV cache. Soft tokens are a learned latent of the dropped history, close in spirit to Context Memorization (05-20). 5.2% of positions is a direct prefill and KV saving, and unlike a rewritten summary it does not have to be re-read from disk.
- Swarm vs compaction. The essay's case for swarms is that a sub-agent's fresh context is cheaper than rebuilding understanding after compaction. Anthropic shipped exactly that this week (Claude Managed Agents dynamic workflows, up to 1,000 sub-agents; one agent found at most 27 of 70 planted bugs, the workflow 66). The multi-agent systems trend now has a capacity argument, not just a parallelism one.
Gaps
- REMORY needs training per base model; transfer across model families is not shown.
- The swarm essay is an argument, not a measurement; Prime Agent's rewrite has no published cost.
Related
Agent memory · Multi-agent systems · KV cache · Agent harness engineering