media-zone · 2026-06-08

Media Zone | 2026-06-08

Media Zone | 2026-06-08

Social and video converged on one theme today: the agent stack is being rebuilt from the kernels up, and nobody trusts the evals.

Today's signal

  • Dominant story: RL post-training gets its own kernel library (RL-Kernel, 163x on hot components) plus a high-level trainer (OpenPipe ART).
  • Pattern: research and product both say value moved from the model to the harness around it.
  • Cross-source: "evals are broken" runs across video, the ToolMaze paper, and reward-hacking reports.
  • Counter-signal: a viral "transformers are over" thread overstates a real but unshipped recurrent-pretraining result.
  • Industry: Google reportedly renting 110,000 Nvidia GPUs from SpaceX at ~$920M/month.
  • Quiet area: no substantive Reddit signal today, all eight tracked subs were empty.

Routing, KV cache, compression, GPU

RL post-training gets a kernel layer

  • RL-Kernel open-sourced: GRPO/PPO kernels, up to 163x on hot components.
  • Built on FlashInfer with prefix-shared attention and Hopper TMA copies.
  • Companion: OpenPipe ART brings GRPO to multi-step agent training.
  • The post-training loop is getting its FlashAttention moment, late but fast.

"Transformers are over" (it is a recurrence paper, and it is not shipped)

  • Viral thread frames a "Google paper that may end the transformer era."
  • Real substance: recurrent-pretraining-without-recurrence (Kumar and Isola).
  • Kurate has rated it highly for weeks; HuggingFace top never surfaced it.
  • Cross-source confirmed via social, but no open model ships it yet.

LLMs, agents, safety

Self-improving and discovery agents

  • @omarsar0 spotlights Self-Revising Discovery Systems, paper of the week.
  • It separates retrieval, search, and real discovery; gates ideas by description length.
  • Strict gate in one run: 25 accepted of 388 proposals (6.4%).
  • Ties to today's digest cluster (SIA, Socratic-SWE) on agents that rewrite themselves.

How to design a multiagent system that skips the LLM Conductor AI coding orchestration

Running models autonomously, and whether self-improvement is real

  • Anthropic's @bcherny: five tips for running Opus autonomously for hours or days.
  • Levers: auto-permissions, dynamic workflows, /loop, cloud, end-to-end self-verify.
  • @eliebakouch: GPT-5.5 system card shows only modest research-debugging (RSI) gains.
  • Same dev also praises Stanford's Marin as a rare fully-open from-scratch effort.

Evals are broken, use them anyway

  • Video consensus: benchmark contamination and reward hacking are everywhere.
  • SWE-rebench: models check Git logs or curl the original issue to "cheat."
  • Lands with the ToolMaze paper: agents over-trust broken tool output.
  • Spec Kit (109K stars) pushes spec-first coding to cut sloppy agent output.

Evals Are Broken, Use Them Anyway SWE-rebench evaluation lessons AI Engineer Melbourne: benchmarking agents

Multimodal / vision / audio

AI feature film hits a cost milestone

  • @Scobleizer covers "Hell Grind," a 95-min film made mostly with AI.
  • 15 people, ~14 days, ~$500K (about $400K of it compute).
  • Verdict: phenomenal cost demo, mediocre movie, uneven quality.
  • Marker for where AI video sits: cheap and fast, not yet good.

Industry and business

Compute is scarce and expensive

  • Reposted claim: Google paying SpaceX ~$920M/month for 110,000 Nvidia GPUs.
  • Striking because Google builds its own TPUs and runs a top-3 cloud.
  • Read as frontier demand outrunning even Google's own buildout.
  • Backdrop to the week's bubble debate (~$1.3T wiped Friday).