media-zone · 2026-07-27

Media Zone | 2026-07-27

Media Zone | 2026-07-27

Social spent the weekend arguing that the workflow around Opus 5 matters more than Opus 5, while Kimi day arrived with a capability fight already in progress.

Today's signal

  • Dominant story: Kimi K3 ships today, and the cyber-capability argument preceded the weights.
  • Pattern: three separate threads all say the harness beats the model.
  • Counter-signal: an eval token cap may be hiding top-tier cyber capability, not measuring it.
  • Price war is the loudest xAI message: Grok 4.6 in two weeks, 4.7 in four.
  • Quiet area: zero curated reposts, zero Reddit posts cleared filters across all eight subs.
  • No new YouTube in the sync since 07-25, so this slot is Twitter-only.

Routing, KV cache, compression, GPU

The workflow, not the model, is the variable

  • minchoi: "most people are using Opus 5 wrong" after a disappointed first pass.
  • Claim: the right agent workflow lets Opus 5 outperform Fable.
  • Same week Cursor split planner from worker and cheap models hit 100%.
  • Social reached the routing conclusion by vibes, papers reached it with numbers.
  • See today's digest on latent-side routing at 90.7% fewer frontier calls.

Write code that a grep can find

  • Modem's codebase: 680K lines, 99.9% LLM-generated, since Sonnet 3.7.
  • Core claim: agents navigate by string search, not by reading modules.
  • Three levers you control: names, types, where explanations live.
  • xAI engineer's framing: always been true, now it has a price tag.

LLMs, agents, safety

Kimi day, and the eval that cannot see K3

  • tinygrad flagged Kimi day for July 27 a week early.
  • Bakouch: K3 may be top-tier ("Mythos") for offensive cyber capability.
  • It does not show in UK AISI's eval because that eval caps at 100M total tokens.
  • His read: K3 is not reasoning-efficient enough to prove itself inside that budget.
  • Independent eval puts K3 between Opus 4.8 and GPT-5.6 Sol.

Open weights at the top tier, argued in public

  • Lebovic: closed models are not solving model-access-for-defenders.
  • So open weights preferable even at Mythos level, if the capability exists at all.
  • Bakouch on a global capability pause: would press the button, would not lobby for it.
  • His line: infeasible globally, feasible nationally, and he opposes the national one.
  • Note the gap with policy, where Washington reportedly favors targeted bans.

Opus 5 demos are still demos, three days in

  • minchoi thread: one-shotting 3D games, worlds, Blender builds.
  • Still no independent evaluation on social, third day running.
  • The one hard number came from press: 30.2% on ARC-AGI-3.
  • That is nearly 4x GPT-5.6 Sol's prior 7.8% record.
  • Reported behavior: independently formulating reflection equations.

Multimodal / vision / audio

Real-time avatars stop taking turns

  • Scoble on Vivix A1: it listens, responds, and moves simultaneously.
  • You can talk over it or change direction mid-sentence without a reset.
  • Facial expression, eye gaze, voice, and body movement stay synchronized.
  • Accepts voice, text, and images at any point, and places objects into the scene.
  • Turn-taking removal is the interesting bit, not the rendering.

Industry and business

xAI competes on cadence and price, not benchmarks

  • Musk: Grok 4.6 in two weeks, Grok 4.7 in four weeks.
  • Super heavy tier at $100, roughly half Codex and Claude 20x plans.
  • Claimed nearly double the rate limits of both.
  • Argil churned from Claude Code to Grok Build, citing speed above all.
  • Tim Sweeney rates Grok 4.5 level with Sol on programming-language-theory work.