Agensh: Scaling Organizational Intelligence to 1,024 Agents
Source: arXiv 2609.26781 (Microsoft Research). Summary: DAIR Academy. Surfaced via the X home feed (@omarsar0), 2026-09-28. Raw: raw/twitter/feed/2026-09-28-morning-ranked.json. No alphaxiv overview yet; based on the thread and abstract summary.
TL;DR
Most multi-agent coding systems have a boss: an orchestrator that splits the work and hands it out. Agensh removes the boss. Parallel coding agents coordinate only through shared state: a shared workspace and a message channel. Each agent gathers context, claims a sub-task, does it, shares what it learned, verifies the result and merges it, all asynchronously. On the five hardest ProgramBench tasks with GPT-5.6-sol, going from 1 to 128 agents raises the mean final test-pass rate from 19.31% to 28.78%, and bigger teams reach a given pass rate sooner. On pandoc, 1,024 agents move the pass rate from 33.89% to 55.06%. The authors report cooperation habits the agents invent on their own that become standard as the team grows.
flowchart LR
W["Shared workspace<br/><small>code, tasks, messages</small>"] --> C["Claim<br/><small>pick an open sub-task</small>"]
C --> K["Work<br/><small>one coding agent</small>"]
K --> S["Share<br/><small>post findings</small>"]
S --> V{"Verify<br/><small>tests pass?</small>"}
V -->|yes| M["Merge<br/><small>into shared code</small>"]
V -->|no| C
M --> W
classDef input fill:#d0ebff,stroke:#1971c2,color:#1b1b1b,stroke-width:2px
classDef core fill:#e5dbff,stroke:#6741d9,color:#1b1b1b,stroke-width:2px
classDef loop fill:#fff3bf,stroke:#f08c00,color:#1b1b1b,stroke-width:2px
classDef exit fill:#d3f9d8,stroke:#2f9e44,color:#1b1b1b,stroke-width:2px
classDef err fill:#ffe3e3,stroke:#e03131,color:#1b1b1b,stroke-width:2px
class W input
class C,K,S core
class V loop
class M exit
linkStyle 4 stroke:#2f9e44,stroke-width:2px
linkStyle 5 stroke:#e03131,stroke-width:2px
Key points
- Scale without hierarchy. Coordination cost is pushed into shared state, not a central planner that would become the bottleneck.
- Returns are real but sublinear. 128x the agents buys about +9.5 points on the hardest tasks; 1,024 agents buy about +21 points on pandoc. Cost per point rises steeply.
- Emergent conventions. The agents develop cooperation practices unprompted, which the paper documents as a finding rather than a design.
Relation to prior wiki pages
- Contrasts with SAT (09-24). Self-Organizing Agent Teams learned reusable team strategies for small, fixed teams on reasoning tasks. Agensh scales team size by three orders of magnitude and gets organization from shared state rather than learned roles. See multi-agent-systems.
- Harness thesis, at the org level. agent-harness-engineering has argued that the loop around the model, not the model, sets outcomes. Agensh is that argument for many loops at once.
Gaps
- Token cost per point of pass rate is the missing number; 1,024 frontier-model agents is not cheap.
- Five tasks plus pandoc is a narrow slice; no comparison against a strong orchestrated baseline at matched spend.