media-zone · 2026-08-01

Media Zone | 2026-08-01

Media Zone | 2026-08-01

A hardware shop out-benchmarked a model vendor, a lab posted proofs instead of a benchmark table, and the loudest AI accounts spent the day on geopolitics.

Today's signal

  • Dominant story: tinygrad's DeepSeek V4-Flash serving table beat the vendor's own runbook number.
  • Pattern: every big post today came with a receipt. Lean proofs, benchmark tables, red-team lists.
  • Counter-signal: three separate posts on agents being unreadable, uncountable, and wasteful to supervise.
  • Quiet area: no new AI YouTube since 07-29, and all eight Reddit farms came back empty.
  • Noise floor: roughly two-thirds of the AI handle feed was geopolitics from non-AI accounts.

Routing, KV cache, compression, GPU

tinygrad beats the runbook, then sells you a smaller box

  • 245 tok/s single user on DeepSeek V4-Flash, only 2 Blackwell GPUs busy.
  • Runbook's own validated number was 217 to 220. Practitioner beat vendor.
  • Stack: W4A8 kernels, fp8 KV cache, DSpark K5 speculative decode, 131k context.
  • The footnote is the story: acceptance 90.5% on synthetic prompts, about 64% on real code.
  • tiny corp shipped a 2-GPU tinybox variant within the hour.

The price collapse became a mood, not an argument

  • OpenAI credits speculative decoding for 20% lower serving cost, then cuts Luna 80%.
  • Social read: Luna matches March's flagship at one-thirteenth the price.
  • @ns123abc's whole take on DeepSeek pricing: "lol deepseek is literally free."
  • NVIDIA counter-programs with a customer case study claiming roughly 10x cheaper inference.
  • Nobody in any thread asks whether the speculative verifier is lossless.

LLMs, agents, safety

Astra: the proofs travelled further than the model

  • @ns123abc's enumerated list, not OpenAI's post, is what circulated.
  • Non-sofic groups, Connes rigidity disproven, three Erdős problems, exact sphere-packing bound.
  • Four of ten results are counterexamples to what mathematicians believed.
  • Lean 4 repository attached, which is why nobody is arguing about validity yet.
  • Mathematician Thomas Bloom's "big news" is doing a lot of the credibility work.

Open weights got a procedure and a number in the same 24 hours

  • Murati amplifies the framing: indiscriminate release unsafe, lab-only capture also unsafe.
  • Four red teams with disjoint mandates, plus fine-tuning that strips safety training.
  • Kilo's counterpart is telemetry: open weights are 79.1% of its token usage.
  • Kilo's line lands harder than the framework: "hiding the number doesn't change the number."
  • Anthropic remains the one big name absent from the open-weights letter.

Agents outran the humans watching them (three posts, three layers)

  • Santiago: "my mind is no longer able to keep up" with agent context.
  • xAI's Jonas Badalic: agents overcompensate with sophisticated-sounding answers, same as humans.
  • His test: ask for an explain-like-I'm-five and get two confidently wrong sentences.
  • DHH ships an Omarchy panel tracking Claude and Codex burn. GPT-5.6 Sol at 165.5M tokens.
  • Min Choi's prompt: bill like a contractor, never ask what the repo already answers.

Multimodal / vision / audio

Character consistency, not resolution, is the video story

  • Grok Imagine 1.5 adds text-to-video, image and voice references, native 1080p.
  • Omni-reference is the actual upgrade: consistent characters across a long story.
  • Amplified by Xiang He of Google DeepMind, a competitor, which is the tell.
  • Nobody posted output. All the enthusiasm is about the conditioning mechanism.

Industry and business

Robots and defence money moved, quietly

  • Xiaomi's factory humanoid: 98% accuracy after four months on one Beijing EV station.
  • Two more task types already above 90%. Four months beats any demo video.
  • US Department of War commits $820M loan to Performance Drone Works manufacturing.
  • Scoble's only comment on the Xiaomi robot: "not available in USA."
  • A Pixel 10 running Gemma 4 locally coached a race car at Sonoma.

Aschenbrenner writes to his LPs

  • Full letter published after Situational Awareness dumped its portfolio to Citadel.
  • Margin calls came days after a reported 439% six-month return.
  • Letter's message: rumours of the fund's demise are exaggerated.
  • Reaction was mostly loyalty posting, not analysis. "I stand with Leopold."