media-zone · 2026-10-06

Media Zone | 2026-10-06

Media Zone | 2026-10-06

The US Monday's conversation has one cost idea underneath it: stop paying frontier prices for work that does not need them. A coding agent rewrote its own harness to Codex level for $4.03, and a Microsoft curriculum method cut harness-tuning spend from $1,360 to $298. Two context papers cut agent tokens by a third to a half while scoring higher. Red Hat shipped a 125B MoE with only its experts squeezed to FP4, and Cloudflare, OpenAI and llm-d all shipped small classifiers that decide where a call goes. GPT-6's 100x price spread makes that routing layer a necessity, not an optimization.

Today's signal

  • Dominant story: harnesses are becoming things you search, train and grade. Four harness papers in one US day: SelfSearch (self-edits, $4.03), ActiveSaddler (pick the failures worth fixing, 4.6x cheaper), a Wavestone teardown of 11 coding agents, and a Harvard-MIT result that "plan ahead" prompts make agents worse.
  • Pattern: context is now a cost line the model manages itself. Context Language Models and CorpusMap both raise accuracy while cutting compute or input tokens by 21-59%.
  • Routing goes mainstream: Cloudflare's open Clef, OpenAI's Decisions API and llm-d's Semantic Classifier are all small models whose only job is to decide where a call goes.
  • Compression in practice: mixed-precision MoE checkpoints (FP4 experts, BF16 everything else) and expert offloading to host RAM keep putting 100B+ models on single GPUs.
  • Counter-signal: memory pricing hit the desktop. NVIDIA's 128GB DGX Spark went from $3,999 at launch to $6,950.
  • Noise: "Karpathy's 24/7 AI employee" threads, "free gold" repo lists and a Stanford orchestration claim citing an unverifiable arXiv id. Skipped.
  • Evening turn: the NYC AI hearing became the US afternoon's loudest thread. A draft bill would fine vendors $25K per unvalidated model sale and require a shutdown capability.
  • Quiet areas: no new X bookmarks in any bookmark run this window (so the saved-reading section is empty today), the curated retweet scrape came back empty, YouTube had no new AI videos, LinkedIn posts arrived without text, and no Reddit post passed the filters. HuggingFace did post a fresh 65-paper list; its KV-cache papers are in the digest, not here, because social did not pick them up.

Routing, KV cache, compression, GPU

Big MoE models, small hardware

The practical compression story of the day: keep the shared parts of a mixture-of-experts (MoE, where each token runs through only a few specialist sub-networks) at full precision, and crush the experts, which hold most of the weights but are each used rarely.

Squeeze the experts, keep the shared path precise
Most MoE weights sit in experts that any one token barely touches, so that is where the bits go.
flowchart LR
  T["Token<br/><small>incoming hidden state</small>"] --> A["Attention<br/><small>BF16, every token</small>"]
  A --> R{"Router<br/><small>picks a few experts</small>"}
  R -->|top-k| E["Experts<br/><small>FP4, mostly idle</small>"]
  E --> O["Output<br/><small>next layer</small>"]
  E -.->|offload| H["Host RAM<br/><small>cold experts parked</small>"]
  classDef input fill:#d0ebff,stroke:#1971c2,color:#1b1b1b,stroke-width:2px
  classDef core fill:#e5dbff,stroke:#6741d9,color:#1b1b1b,stroke-width:2px
  classDef loop fill:#fff3bf,stroke:#f08c00,color:#1b1b1b,stroke-width:2px
  classDef exit fill:#d3f9d8,stroke:#2f9e44,color:#1b1b1b,stroke-width:2px
  classDef err fill:#ffe3e3,stroke:#e03131,color:#1b1b1b,stroke-width:2px
  class T input
  class A core
  class R loop
  class E err
  class H err
  class O exit
  linkStyle 2 stroke:#f08c00,stroke-width:2px
Blue is input, purple runs at full precision, amber is the routing decision, red is the compressed or offloaded part, green is the result.
Hugging Face · quantized checkpoint

Qwen3.8-Flash-Next in NVFP4, experts only

Red Hat AI published an NVFP4 (NVIDIA's 4-bit floating-point format with fine-grained scaling, native on Blackwell) checkpoint of Qwen3.8-Flash-Next. Only the MoE experts are quantized to FP4. Attention, embeddings and the router stay in BF16. The model takes text, image and video and loads directly in vLLM. Red Hat posted benchmarks against other popular checkpoints, and a team member separately pitched their Hub as the place for "the most accurate quantized checkpoints." This was the top-ranked post in the feed for fit and engagement rate. Expert-only quantization is the cheap, low-risk default: it cuts most of the memory while leaving the precision-sensitive shared path alone.

Newsletter · expert offloading

The same 125B model on a 12GB gaming GPU

AI Weekly's toolbox picked up Strata, an open-source engine that claims to run the same 125B-parameter Qwen3.8-Flash-Next on one 12GB consumer GPU with 32GB of system RAM. It parks experts in host memory and pulls them in as the router needs them. The project quotes roughly 100 to 140 tokens per second on an RTX 3090. Those are the project's own numbers, so wait for LocalLLaMA reproductions. Together with the NVFP4 release, it shows two routes to the same goal: shrink the experts, or move them off the GPU.

via AI Weekly Espresso (10-05)
Hardware · memory pricing

DGX Spark gets a 64GB model, and the 128GB one jumps 74%

NVIDIA added a 64GB DGX Spark at $4,999, shipping October 23, on the same GB10 superchip with half the unified memory. Two units cluster to 128GB over ConnectX-7. The other half of the news: the 128GB model is now $6,950, up from $3,999 at launch. AI Breakfast read it as memory shortage pricing reaching a desktop box rather than a datacenter invoice. It also explains the demand for expert offloading and 4-bit checkpoints above.

Context is a cost line: let the model manage it

Two papers from the US afternoon attack the same bill from different ends: the tokens an agent re-reads. One makes the model edit its own context. The other pre-builds a map so the agent stops re-searching.

Paper · UW, Meta Superintelligence Labs, MIT · context management

Context Language Models: the context is a file the model edits

Today most agents rely on the harness to decide what stays in the context window, with fixed rules like "summarize after N tokens." Context Language Models (CLMs) hand that job to the model. The context is treated as a file, and the model can rewrite it freely: compress, delete, keep or reorganize. Built zero-shot from existing models, CLMs scored 11.4% higher on BrowseComp-Plus (a deep web-research benchmark) with 21.5% fewer FLOPs, and 5% higher on a 12-hour EdgeBench run with 59% fewer FLOPs. Because the policy now lives in the model, RL can train it, and the authors show it improves further. The author list (Zettlemoyer, Lewis, Yih, Lambert, Koh) makes this one to read. It is the learned answer to the UT Austin context-compression study from 10-05, which compared hand-built compression rules.

Paper · Microsoft · agentic search

CorpusMap: follow the entities, stop re-searching

When an agent searches a large document set exposed as a flat folder, a useful document gives no hint about which other documents relate to it. So the agent rediscovers the same links on every query, misses evidence and burns tokens. CorpusMap builds, offline and without LLM calls, an "entity page" for every recurring person, project or product, linking every document that mentions it. The agent reads a document, then follows an entity to the related ones. Across 7 models and 3 benchmarks, answer quality rose 6.4 to 11.7 points while input tokens fell 34% to 57%. It also beat an LLM-written wiki layer. That is close to how this wiki itself works, and it shows the cheap, deterministic version wins.

  • Leviathan, from the overnight US capture: an open-source single-binary indexer that turns JSONL, CSV or SQLite records into a ranked full-text index for agents. The author claims that at 1M records it hands the agent 436 tokens instead of 107,000 and finds the answer 99% of the time. It is the cheapest version of the same idea as CorpusMap: structure once, read little (X · repo).
  • Chunked prefill, explained: Akshay Pachaar's walkthrough of how vLLM splits a long prompt's prefill into chunks so it does not stall other requests' decode steps was the top post of the morning capture. A good refresher on why prefill-heavy agent traffic needs scheduler work, not just more KV memory (X).

Decoding faster: smarter drafts, or no left-to-right at all

vLLM · speculative decoding

Variable-length speculation without a confidence head

Speculative decoding has a cheap draft model guess several tokens ahead. The big model then checks them in one pass. Deciding how many tokens to draft used to need an extra trained confidence head. vLLM 0.30 generalized "adaptive verification" to any draft-based method, including Eagle and DFlash, using an online acceptance estimator that learns the acceptance rate as it serves. Draft length now adapts to how predictable the text is, with no retraining. Version 0.31 is already out and will be covered at the next vLLM office hours.

DeepMind · diffusion language model

Write 256 tokens at once, then refine

A widely shared breakdown says Google DeepMind shipped a language model that drafts a 256-token block in parallel and refines the whole block repeatedly, instead of writing one token at a time. It is described as a 26B MoE with 3.8B active parameters. The pitch is editing work: infilling, formatting fixes and keeping several sections consistent, where seeing the whole output helps. DeepMind's Sander Dieleman called 2026 "the year of continuous diffusion for language" and re-shared his long post on why continuous diffusion, written off a few years ago, is coming back. For serving cost, parallel decoding swaps memory-bound token-by-token decode for fewer, more compute-heavy passes. Those are the same economics speculative decoding is chasing.

Decision models: the router becomes a product

vLLM / llm-d · routing classifier

llm-d Semantic Classifier, and why it is not a guardrail

The vLLM Office Hours #58 recording went up this afternoon. Beyond the v0.29 and v0.30 changes and a new Tenstorrent hardware backend (one more non-NVIDIA chip that vLLM can serve on), it introduced the llm-d Semantic Classifier. llm-d is the Kubernetes-native distributed serving layer built around vLLM. The classifier is a lightweight service that labels requests at high throughput so the serving layer can route them. The talk explicitly argues it is not a guardrail: it exists to send traffic to the right model or pool, not to block it. With Clef and the Decisions API below, that makes three "small model decides where the call goes" launches in one day, one of them inside the open serving stack.

Cloudflare + OpenAI · gating layer

Clef and the Decisions API return a probability, not prose

"Decision models" take some input and a fixed list of possible answers, then return a calibrated confidence for each. They are built for routing an email, scoring a lead or choosing an agent's next step. Jev started the category. Cloudflare's open-weight Clef (27B) and Clef-flash (9B) are Apache 2.0, read text, JSON, images and video across 65K tokens, and cost $0.24 and $0.09 per million input tokens on Workers AI. Cloudflare reports them 2.5x and 13x faster than Jev at the median, behind a Jev-compatible API. OpenAI's Decisions API, built on GPT-6 Luna, is in limited preview. This is the routing layer from the 10-05 cache-aware routing post, sold as a product. The open weights matter because a gating model sits on every call, so self-hosting it is where the savings are.

Pricing · routing pressure

GPT-6 spans 100x, and the fast tier burns 8x

OpenAI's GPT-6 build guide lists Astra at $10/$50, Sol at $2/$10 and Luna at $0.10/$0.50 per million input/output tokens. That is a 100x spread inside one model family. OpenAI also launched a $500/month Pro tier with "Astra Ultrafast" at up to 8x speed, which uses up the plan's allowance 8x faster. Meanwhile, existing $200 Pro users lose part of their allowance after October 29. A feed thread made the same point from the local side: an agent makes a thousand calls an hour, and maybe forty need a frontier model. Small accounts also pushed "progressive determinism": once an agent's query has been verified, cache it as code so it never goes back to the LLM.

Chips, power and who builds the fabs

  • TSMC may help run Musk's Terafab fabs in Texas, per a report Musk called "just discussions." TSMC rose and Intel fell premarket, though Intel can still win 14A orders there (X).
  • Qualcomm will pay Huawei under a patent deal for the first time. It covers 5G, AI, 3D stacking and near-packaged optics, the layers where scaling now happens (X).
  • Power is the scarce asset. TeraWulf doubled contracted power at one site to 1 GW, and its Anthropic lease implies about $2.4B a year of revenue at that scale (X). Toshiba will double datacenter hard-drive output by FY2027 to fill the AI memory gap (Nikkei).
  • NetworkOcean (YC) floated an H100 at sea: OASIS-00 runs on floating solar with seawater cooling, sketch to water in 53 days. A stunt today, but it shows where power-hunting is heading (X).
  • PyTorch's Accelerator Integration Working Group published its yearly update: a cross-repo CI relay, device-agnostic tests and the OpenReg reference backend, all to make new non-NVIDIA chips cheaper to support (PyTorch).

LLMs, agents, safety

Harnesses become trainable objects

The harness (the loop, tools, prompts and memory wrapped around a model) was the most-discussed theme of the US morning, with two concrete artifacts behind it.

The agent improves the agent, with no task reward
SelfSearch learns from records of earlier self-edits, not from benchmark scores.
flowchart LR
  A["Agent v0<br/><small>prompts, tools, procedures</small>"] --> M["Self-edit<br/><small>rewrites its own harness</small>"]
  L["Edit records<br/><small>reasoning, actions, outcomes</small>"] --> M
  M --> N["Agent v1<br/><small>becomes next improver</small>"]
  N -->|logs attempt| L
  N --> B["Benchmark<br/><small>scored only at the end</small>"]
  classDef input fill:#d0ebff,stroke:#1971c2,color:#1b1b1b,stroke-width:2px
  classDef core fill:#e5dbff,stroke:#6741d9,color:#1b1b1b,stroke-width:2px
  classDef loop fill:#fff3bf,stroke:#f08c00,color:#1b1b1b,stroke-width:2px
  classDef exit fill:#d3f9d8,stroke:#2f9e44,color:#1b1b1b,stroke-width:2px
  class A,N core
  class L input
  class M loop
  class B exit
  linkStyle 3 stroke:#f08c00,stroke-width:2px
Purple is the agent, blue is the experience log, amber is the self-modification loop, green is the final evaluation.
Paper · Seoul National University · self-improving agents

SelfSearch: a Codex-level harness for $4.03

Coding agents can read and edit their own instructions, tools and procedures. Earlier work searched over those edits by re-running benchmarks after each change, which is expensive and overfits to the benchmark. SelfSearch drops the reward. Each agent edits itself using records of earlier edit attempts (the reasoning, tool actions and what happened), and the edited agent becomes the next improver. Average success rose over the starting agent in all six model-benchmark settings, by up to 11.2 points on Terminal-Bench 2.1. On SWE-bench Multilingual one evolved agent gained 5.0 points and spent 38.5% less on the tasks both versions solved. With $4.03 of search spend, it produced a harness where DeepSeek V4 Flash solves 82.0% of Terminal-Bench 2.1, matching Codex, the top harness in a public nine-harness comparison. This extends the 10-05 "spend compute on the verifier" thread: the cheapest gains are now in the scaffold, not the weights.

Hugging Face · multi-harness RL

Claude Code, Codex and friends as RL environments

Clem Delangue announced that Hugging Face turned Claude Code, Codex, Hermes, Pi, opencode and other coding harnesses into RL (reinforcement learning) environments with no changes to the harnesses. One open-weight model can now be trained against the real loops it will run in. Akshay Pachaar's guide to multi-harness RL, reshared by Hugging Face, was among the fastest-rising posts this afternoon. This closes the gap where models were trained in one scaffold and deployed in another. Expect "harness-robust" to become a training target.

Paper · Microsoft · harness optimization

ActiveSaddler: tune the harness on the failures still open

Automatic harness tuners improve an agent's prompts, tools and control logic from execution feedback. They focus on how to patch the harness and feed it a fixed task list. ActiveSaddler argues that which tasks produce the feedback matters as much, and that the best tasks change as the harness improves. It groups recurring failures into "failure-pattern arms" and treats picking the next one as a bandit problem (balancing known weaknesses against exploring new ones). With the same optimizer, test pass rates rose 4.4 points on GAIA2 and 7.5 on Terminal-Bench 2.0. Reaching 58.5% on GAIA2 dev cost $298 versus $1,360 with a fixed task order. Paired with SelfSearch, the lesson is that harness search is now cheap enough that its budget allocation is the research question.

Paper · Wavestone AI Lab · harness survey

Eleven coding agents, one seven-part harness

This survey takes apart Claude Code, Codex, Gemini CLI, Aider and seven more, starting from "agent = model + harness." Every one has the same seven parts, differing only in size: the loop (plain while-loop up to replayable runs), the LLM layer, tools (bash only up to 43 typed tools), memory, safety (a step limit up to rules plus a reviewer model plus a sandbox), orchestration (sub-agents or not), and extensions (SKILL.md in 9 of 11, MCP in 8). It notes almost nobody uses agent frameworks or embeddings. It ends with 18 design rules and a 90-line reference harness. It crossed the feed twice, from Spanish- and English-language accounts, which is a fair sign it is becoming the standard reference.

Paper · Harvard and MIT · scaffolds vs interfaces

"Plan ahead" prompts make agents worse. Simpler interfaces fix it.

The authors use auctions and matching markets, where the optimal strategy is known, so every choice can be scored. Prompts that tell the agent to plan through rounds or model the other players made play worse overall. Changing the interface helped instead: an ascending auction, which offers one safe choice at a time, cut bid errors across four model families. Spelling out payoffs and why truthful bidding is safe also helped. The agents' written plans did not track their actual choices. Some changes improved bids without improving the plans, and some improved the plans without improving bids. The practical rule: judge a scaffold by the decisions it produces, not by how good its reasoning reads.

Debate · who owns the harness

Lab harnesses burn tokens. Verification is the new bottleneck.

Garry Tan argued that lab-built harnesses have an incentive to burn tokens, and Y Combinator amplified it into the evening. Startup harnesses are useful because they can watch agent behavior and replace repeated token spend with deterministic, tested code. Elvis Saravia (omarsar0) says he now spends more time verifying agent-written code than writing it. ContextQA's Ship, an autonomous QA agent that reproduces Slack or Linear bug reports and hands the context to Claude Code or Codex, launched into that gap. NVIDIA open-sourced 391 signed agent skills with benchmarks (CUDA, Jetson, RAG, robotics) under Apache 2.0. Anthropic's Claude Mods, which let users change how Claude Code itself works, got its first follow-up releases.

Agents meet the web and the enterprise

  • Six ways an agent reaches a website, from raw API to WebMCP (which Chrome and Edge are building together). Pages declare their actions with names and typed inputs, so agents stop paying vision tokens to find a button (thread).
  • Personal agents vs "factories." Nathan Baschez argues the end state is shared multi-agent systems, because they allow specialization and trade. ChatGPT Dots, Grok bots and Muse are the personal-agent wave he is arguing against (essay).
  • Cohere North 2 adds reusable agents, multi-agent orchestration and memory across sessions, pitched on sovereignty and security (blog).

Grade the outcome, not the agent's last message

The evening's agent-eval theme: stop trusting what the agent says it did, and check the state it left behind.

  • Microsoft ThinkingBox is now an OpenEnv environment. It simulates a customer, gives the agent MCP tools over a real database, then grades what changed in the database, and asks whether the agent can do it twenty times in a row. ThinkingBox-Bench has 507 tasks across retail, insurance, travel, banking and consulting (HF blog · X).
  • Era by Eon generates a whole simulated company across Salesforce, Zendesk, Slack, Jira and Gong, down to vendor rate limits and error codes. Because Era wrote every record, code computes each answer exactly, so grading needs no LLM judge. NVIDIA is a research partner (Era · omarsar0 · rohanpaul).
  • flow-1, a YC-backed model trained with RL to find errors in agent traces, claims to match GPT-6 Sol at trace analysis. Only the launch tweet so far (RT).
  • Lizard moved its agent sandboxes to Firecracker microVMs (one kernel per sandbox) with templates that ship Claude Code, Codex, opencode and Pi preinstalled, pitched as the cheapest on the market. Vendor post, but the highest engagement rate of the evening (X).

Swarms: speed, or new capability?

  • Understanding AI asks whether agent swarms are the next scaling law. Noam Brown credits under 10% of OpenAI's Navier-Stokes result to the 10,000-agent swarm. The Claude Opus 5.5 system card found most multi-agent gain comes from going 1 to 10 agents. Toby Ord estimates a "stepping on toes" parallelism factor of 0.5 to 0.7 for GPT-5.6 Sol swarms, close to human teams (essay).
  • The counterpoint is a Microsoft-Berkeley paper finding that bigger coding-agent teams score higher than solo agents even with generous time, and that one task was only solvable by a team. It resurfaced on X this evening (RT). For cost, the open question is whether swarms buy capability or just buy wall-clock time at a token premium.

Safety and policy, live from New York

  • NYC City Council AI-risk hearing ran through the US day. Gary Marcus testified and pushed back on ex-Anthropic researcher Jacob Coxon's "we don't know how to control any AI system." Marcus says that is true of generative AI specifically, so pause that (X · testimony).
  • The bill behind the hearing: no AI model sold or deployed in NYC without outside validation (accuracy, calibration, determinism, latency, provenance, bias, privacy, safety), certified to the city's Cyber Command, plus a mandatory shutdown capability and a $25K fine per unvalidated sale (X).
  • Testimony highlights: ex-DeepMind researcher Alex Turner described Google backing away from earlier safety commitments. Daniel Kokotajlo said the ability to notice misalignment is "quite poor and set to get much worse," and argued for several independent evaluators. Speaker Menin asked firms to testify under oath that their systems follow safety instructions (Turner · Kokotajlo · Menin).
  • OpenAI will watermark text in the EU to meet regulation, while saying plainly that text watermarks are often undetectable in short passages and vanish under rewriting or translation. The detector goes only to approved researchers for now (X · OpenAI).
  • Trump's "Super Intelligence Force," chaired by DNI Jay Clayton, has 120 days to report, with a charter warning against overregulation (TechCrunch). Researchers on X mocked the forced "SI" relabeling.
  • Anthropic's human reviewers reported a Florida user's threatening Claude chat to police, at least the third such referral since August, per Tom's Hardware (link).
  • Mustafa Suleyman called Anthropic's treatment of Claude as possibly conscious "a dangerous anthropomorphism." Pedro Domingos argued unfaithful chains of thought are hallucinations, not deception (X · Domingos).

Multimodal / vision / audio

  • Deepgram Flux TTS keeps acoustic state across a streaming conversation, so the voice keeps its tone after interruptions. It tracks which words the caller actually heard, and first audio arrives in under 200 ms (X · product).
  • NVIDIA open-sourced an image-to-explorable-3D-world tool. It was among the fastest-rising Hugging Face reshares, but it is off-lens for this reader (RT).

Industry and business

  • OpenAI will test visual ads during ChatGPT image generation in the US later this month. It cites 1.2B weekly users and a WeightWatchers CPA 15.3% below paid search (OpenAI).

  • Beijing "transfer stations" resell Claude at 70-90% off through overseas accounts, sometimes passing off Qwen output as Claude (The Information, via AI Weekly Espresso).

  • US-China model gap at a record-low ~3%. DeepSeek V4.1 Flash scores 81.1 on LiveBench against Anthropic's 83.4 (Bloomberg).

  • OpenAI and Synopsys are co-developing GPT-Synopsys for chip-design workflows (The Stack).

  • Microsoft and Meta are rationing Claude. Microsoft cut projected internal Claude spend by over a third and capped its cloud division at $10,000 per employee a month; Meta halved Claude Code seats to 30,000 (The Decoder).

  • SemiAnalysis priced the plans: Anthropic's mid-tier plan is worth ~5x OpenAI's in API terms, and subscriptions use 40%+ of Anthropic's compute for ~10% of revenue (SemiAnalysis).

  • Reka's Rho-1 (19B, text, image, video and robot actions in one context) was pitched on X as "one state, one loop, no tool calls," a single-network alternative to routing between specialist models (X · The Decoder).

Funding, valuations, and compute deals

  • Ex-Groq engineers sued Groq's board over the $20B "non-exclusive" NVIDIA deal, alleging common shareholders were short-changed. The DoJ is already probing it (FT, via AI Weekly Espresso).
  • Schneider Electric is near a $20B+ deal for PTC (Bloomberg). Firmus (NVIDIA-backed) is reserving half of its A$5.5B ASX IPO for existing holders (Bloomberg).
  • Anthropic committed $100M to train 10,000 "Frontier Deployed Engineers" by end of 2027 (Anthropic). Reflection (NVIDIA-backed) is readying its first open-weight model (Axios).
  • BMO initiated Vertiv at Outperform: about 85% of revenue is datacenter, with roughly $3.5M of content per MW that should rise as liquid cooling spreads (X). Higgsfield ranks #5 on a16z's consumer AI revenue leaderboard (RT).
  • Melius upgraded Microsoft on the thesis that enterprises want a "secure wrapper" that routes models and governs agents, which is a sell-side note making the routing argument (X).

Also crossed your feeds

Overnight US capture: Tengyu Ma argues GRPO-style RL cannot be optimal and that self-play and RSI need better RL methods (RT) · an X article on training choice-order invariance into JEV-style decision models, a practitioner's notes from building System One models (X) · "personal agents give your company a split brain," on agents acting on stale private copies of company state (X) · Yann LeCun at ETH Zurich telling academics not to work on LLMs (RT) · an evals-first argument for becoming AI-native (X) · a Peking University system that rediscovered Newton's laws from coordinates alone (single-source thread) (X). New this evening: Episteme ran four models as forecasters on 25 open scientific hypotheses; the interesting part is where they disagree (post) · Gautam Kamath wants AI to find simpler proofs of known theorems, and Caltech's Mathathon is built on exactly that (thread · Mathathon) · the Nobel medicine committee declined to link this year's neuron-switching prize to neural networks (X) · LinkedIn: data annotators lead its estimate of 750K+ new US AI jobs, at a $51K median (RT) · a single-source claim that Claude Cowork tasks move to Anthropic's servers tomorrow; check the official changelog before acting on it (X) · Atlassian's AI SDLC playbook (blog) · "How to keep learning in the age of LLMs" (essay). Resurfaced from 10-05: Google's VeriHarness (agreement across rollouts can hide shared errors) and CheatBench's "Don't cheat!" result (X) · the UT Austin context-compression study (RT) · Schmidhuber on the 1991 linear Transformer, whose cost scales linearly with input length (link) · PyImageSearch's DeepSeek-V3-from-scratch tutorial (link) · a 2M-citation study finding that answer engines cite pages matching the prompt, with content in the top third, and SEO tricks barely matter (X) · Meta's six Muse Spark math papers (Meta) · Mercor: Claude Opus 5 hit 100% on month-end close where CPAs averaged ~37% (Mercor) · GitLab patched a CVSS 9.9 AI Gateway flaw (THN) · Suleyman: $100B training runs are coming (X) · ex-OpenAI safety writer David Robinson's "culture is broken" essay (TechCrunch) · Nollahealth's AI can now prescribe acne treatment in Utah (X) · Cambridge opened a PhD in AI and Society (link). Skipped: a "DeepGEMM saves 98% on tokens" thread (DeepGEMM is DeepSeek's FP8 matrix-multiply kernel library; it speeds up GPU math, it does not cut token counts), "Boris runs 100+ agents in self-improving graphs" engagement bait, "50 repos replace your subscriptions" and "300+ agents leaked" lists, keep4o campaign posts, the "Stanford orchestrates Opus and Astra" thread (single-source claim, unverified), Karpathy course and "AI employee" bait, second-brain repo hype, stock-picking posts, robot-gadget reposts and politics retweets.