Media Zone | 2026-07-27
Social spent the weekend arguing that the workflow around Opus 5 matters more than Opus 5, while Kimi day arrived with a capability fight already in progress.
Today's signal
- Dominant story: Kimi K3 ships today, and the cyber-capability argument preceded the weights.
- Pattern: three separate threads all say the harness beats the model.
- Counter-signal: an eval token cap may be hiding top-tier cyber capability, not measuring it.
- Price war is the loudest xAI message: Grok 4.6 in two weeks, 4.7 in four.
- Quiet area: zero curated reposts, zero Reddit posts cleared filters across all eight subs.
- No new YouTube in the sync since 07-25, so this slot is Twitter-only.
Routing, KV cache, compression, GPU
The workflow, not the model, is the variable
- minchoi: "most people are using Opus 5 wrong" after a disappointed first pass.
- Claim: the right agent workflow lets Opus 5 outperform Fable.
- Same week Cursor split planner from worker and cheap models hit 100%.
- Social reached the routing conclusion by vibes, papers reached it with numbers.
- See today's digest on latent-side routing at 90.7% fewer frontier calls.
Write code that a grep can find
- Modem's codebase: 680K lines, 99.9% LLM-generated, since Sonnet 3.7.
- Core claim: agents navigate by string search, not by reading modules.
- Three levers you control: names, types, where explanations live.
- xAI engineer's framing: always been true, now it has a price tag.
LLMs, agents, safety
Kimi day, and the eval that cannot see K3
- tinygrad flagged Kimi day for July 27 a week early.
- Bakouch: K3 may be top-tier ("Mythos") for offensive cyber capability.
- It does not show in UK AISI's eval because that eval caps at 100M total tokens.
- His read: K3 is not reasoning-efficient enough to prove itself inside that budget.
- Independent eval puts K3 between Opus 4.8 and GPT-5.6 Sol.
Open weights at the top tier, argued in public
- Lebovic: closed models are not solving model-access-for-defenders.
- So open weights preferable even at Mythos level, if the capability exists at all.
- Bakouch on a global capability pause: would press the button, would not lobby for it.
- His line: infeasible globally, feasible nationally, and he opposes the national one.
- Note the gap with policy, where Washington reportedly favors targeted bans.
Opus 5 demos are still demos, three days in
- minchoi thread: one-shotting 3D games, worlds, Blender builds.
- Still no independent evaluation on social, third day running.
- The one hard number came from press: 30.2% on ARC-AGI-3.
- That is nearly 4x GPT-5.6 Sol's prior 7.8% record.
- Reported behavior: independently formulating reflection equations.
Multimodal / vision / audio
Real-time avatars stop taking turns
- Scoble on Vivix A1: it listens, responds, and moves simultaneously.
- You can talk over it or change direction mid-sentence without a reset.
- Facial expression, eye gaze, voice, and body movement stay synchronized.
- Accepts voice, text, and images at any point, and places objects into the scene.
- Turn-taking removal is the interesting bit, not the rendering.
Industry and business
xAI competes on cadence and price, not benchmarks
- Musk: Grok 4.6 in two weeks, Grok 4.7 in four weeks.
- Super heavy tier at $100, roughly half Codex and Claude 20x plans.
- Claimed nearly double the rate limits of both.
- Argil churned from Claude Code to Grok Build, citing speed above all.
- Tim Sweeney rates Grok 4.5 level with Sol on programming-language-theory work.