media-zone · 2026-07-13

Media Zone | 2026-07-13

Media Zone | 2026-07-13

A quiet Monday. The signal is a benchmark (Long-Horizon-Terminal-Bench) and a compression paper (Requential Coding), not social chatter. Reddit empty, Twitter mostly off-topic.

Today's signal

  • Dominant story: Long-Horizon-Terminal-Bench tops HuggingFace, dense-scored long-task agent eval.
  • Pattern: agent-benchmark work converging on long-horizon and dense scoring.
  • Counter-signal: no harness-normalized comparison yet, so attribution stays murky.
  • Quiet area: no product launches, empty Reddit, no substantive AI Twitter.

Routing, KV cache, compression, GPU

Requential Coding: data-free compression

  • Model compression with self-generated training data, no teacher, no original data.
  • From Andrew Gordon Wilson's NYU group, Kurate rating 8.5/10.
  • A new axis alongside quantization, pruning, distillation.

LLMs, agents, safety

Long-Horizon-Terminal-Bench

  • Dense per-step scoring on long terminal tasks, locates where agents drift.
  • Complements CEO-Bench (long-horizon business sim) and Always-On Agents survey.
  • Raises the model-vs-harness attribution question sharpened by The Harness Effect.