Media Zone | 2026-08-01
A hardware shop out-benchmarked a model vendor, a lab posted proofs instead of a benchmark table, and the loudest AI accounts spent the day on geopolitics.
Today's signal
- Dominant story: tinygrad's DeepSeek V4-Flash serving table beat the vendor's own runbook number.
- Pattern: every big post today came with a receipt. Lean proofs, benchmark tables, red-team lists.
- Counter-signal: three separate posts on agents being unreadable, uncountable, and wasteful to supervise.
- Quiet area: no new AI YouTube since 07-29, and all eight Reddit farms came back empty.
- Noise floor: roughly two-thirds of the AI handle feed was geopolitics from non-AI accounts.
Routing, KV cache, compression, GPU
tinygrad beats the runbook, then sells you a smaller box
- 245 tok/s single user on DeepSeek V4-Flash, only 2 Blackwell GPUs busy.
- Runbook's own validated number was 217 to 220. Practitioner beat vendor.
- Stack: W4A8 kernels, fp8 KV cache, DSpark K5 speculative decode, 131k context.
- The footnote is the story: acceptance 90.5% on synthetic prompts, about 64% on real code.
- tiny corp shipped a 2-GPU tinybox variant within the hour.
The price collapse became a mood, not an argument
- OpenAI credits speculative decoding for 20% lower serving cost, then cuts Luna 80%.
- Social read: Luna matches March's flagship at one-thirteenth the price.
- @ns123abc's whole take on DeepSeek pricing: "lol deepseek is literally free."
- NVIDIA counter-programs with a customer case study claiming roughly 10x cheaper inference.
- Nobody in any thread asks whether the speculative verifier is lossless.
LLMs, agents, safety
Astra: the proofs travelled further than the model
- @ns123abc's enumerated list, not OpenAI's post, is what circulated.
- Non-sofic groups, Connes rigidity disproven, three Erdős problems, exact sphere-packing bound.
- Four of ten results are counterexamples to what mathematicians believed.
- Lean 4 repository attached, which is why nobody is arguing about validity yet.
- Mathematician Thomas Bloom's "big news" is doing a lot of the credibility work.
Open weights got a procedure and a number in the same 24 hours
- Murati amplifies the framing: indiscriminate release unsafe, lab-only capture also unsafe.
- Four red teams with disjoint mandates, plus fine-tuning that strips safety training.
- Kilo's counterpart is telemetry: open weights are 79.1% of its token usage.
- Kilo's line lands harder than the framework: "hiding the number doesn't change the number."
- Anthropic remains the one big name absent from the open-weights letter.
Agents outran the humans watching them (three posts, three layers)
- Santiago: "my mind is no longer able to keep up" with agent context.
- xAI's Jonas Badalic: agents overcompensate with sophisticated-sounding answers, same as humans.
- His test: ask for an explain-like-I'm-five and get two confidently wrong sentences.
- DHH ships an Omarchy panel tracking Claude and Codex burn. GPT-5.6 Sol at 165.5M tokens.
- Min Choi's prompt: bill like a contractor, never ask what the repo already answers.
Multimodal / vision / audio
Character consistency, not resolution, is the video story
- Grok Imagine 1.5 adds text-to-video, image and voice references, native 1080p.
- Omni-reference is the actual upgrade: consistent characters across a long story.
- Amplified by Xiang He of Google DeepMind, a competitor, which is the tell.
- Nobody posted output. All the enthusiasm is about the conditioning mechanism.
Industry and business
Robots and defence money moved, quietly
- Xiaomi's factory humanoid: 98% accuracy after four months on one Beijing EV station.
- Two more task types already above 90%. Four months beats any demo video.
- US Department of War commits $820M loan to Performance Drone Works manufacturing.
- Scoble's only comment on the Xiaomi robot: "not available in USA."
- A Pixel 10 running Gemma 4 locally coached a race car at Sonoma.
Aschenbrenner writes to his LPs
- Full letter published after Situational Awareness dumped its portfolio to Citadel.
- Margin calls came days after a reported 439% six-month return.
- Letter's message: rumours of the fund's demise are exaggerated.
- Reaction was mostly loyalty posting, not analysis. "I stand with Leopold."