ai-industry · 2026-10-08 · Tier 3

Cheap-tier reprice, routing to the PC, and the math drop gets graded (2026-10-08)

Cheap-tier reprice, routing to the PC, and the math drop gets graded (2026-10-08)

Sources: Simon Willison on Haiku 5.5, The Decoder on Haiku 5.5, The Decoder on GPT-6 Intelligent UI, The Information on Microsoft PCs, The Information: "Where is all the automation?", Marcus on AI, Hacker News Digest #320, X Following feed.

TL;DR. The US Wednesday was a pricing and placement day. Anthropic's Claude Haiku 5.5 matched GPT-6 Luna's $0.10/$0.50 per million tokens up to 100K tokens (5x above that), with a tokenizer that uses ~1.25x more tokens than Haiku 4.5, a big OSWorld jump (15.7% to 72.4%), half-price Sonnet 5.5 cache reads, and monthly API credits for Max and Team subscribers. Microsoft moved routine Copilot coding onto a 3-bit on-device model on Nvidia RTX Spark PCs. OpenAI rolled GPT-6 with "Intelligent UI" to all ChatGPT users and says answers can start while the model is still thinking, cutting waits 44%. The OpenAI math release moved from hype to grading: Gary Marcus says the procedure is undisclosed; the Association for Human Mathematics urged mathematicians to stop working with OpenAI; and a new paper argues Lean verification of auto-formalized proofs does not guarantee the natural-language proof is right.

Key points

  • Haiku 5.5 economics. Same list price as Luna under 100K tokens and better benchmarks; above 100K, Luna ($0.20/$0.75 past 272K) is much cheaper. Hidden 1.25x tokenizer cost. Reasoning cannot be disabled; default effort is medium.
  • Microsoft on-device. MAI-Code-1.1-Flash (137B/6.8B active, 3-bit, 256K) replaces Claude Haiku 4.5 for routine Copilot work, at no inference charge. Microsoft Execution Containers (agent sandbox) GA on Windows 11. DeepSeek V4 Flash shown at 1.6 bits (~60GB) on the same PCs.
  • Automation gap. TypeSafe's Diogo Almeida: "100% of LLMs today are optimized for assistance with RLHF," so they fail at unattended automation. Anthropic's Cat Wu describes work stuck at "Claude does 80%, I do 20%." Cognition is now valued at $48B.
  • Math drop grading. "Navier-Stokes lost in translation" (arXiv 2610.08144) argues Lean-verified auto-formalization can still mistranslate the statement being proved.

How this relates to prior wiki pages

Related

Compute economics · LLM routing · Responsible AI