Media Zone | 2026-07-13
A quiet Monday. The signal is a benchmark (Long-Horizon-Terminal-Bench) and a compression paper (Requential Coding), not social chatter. Reddit empty, Twitter mostly off-topic.
Today's signal
- Dominant story: Long-Horizon-Terminal-Bench tops HuggingFace, dense-scored long-task agent eval.
- Pattern: agent-benchmark work converging on long-horizon and dense scoring.
- Counter-signal: no harness-normalized comparison yet, so attribution stays murky.
- Quiet area: no product launches, empty Reddit, no substantive AI Twitter.
Routing, KV cache, compression, GPU
Requential Coding: data-free compression
- Model compression with self-generated training data, no teacher, no original data.
- From Andrew Gordon Wilson's NYU group, Kurate rating 8.5/10.
- A new axis alongside quantization, pruning, distillation.
LLMs, agents, safety
Long-Horizon-Terminal-Bench
- Dense per-step scoring on long terminal tasks, locates where agents drift.
- Complements CEO-Bench (long-horizon business sim) and Always-On Agents survey.
- Raises the model-vs-harness attribution question sharpened by The Harness Effect.