Summary
The afternoon belongs almost entirely to Jev, the decision-only model from TypeSafe AI that emits structured choices and probabilities without ever generating a sentence: eighteen of the slot's seventy posts are about it, and unlike yesterday's launch-day noise this batch contains actual mechanism and actual numbers. The single most useful post in the slot is @anderslie's reverse-engineering of why it is fast, which argues the speed is an inference technique rather than a training result and that you could build a Jev-like API on any open-weight model by forking a cached prefill across N questions and reading constrained logits in parallel. Against that, @nutlope classified 1,018 papers for $0.08 at 256ms median latency, which is the first independent cost measurement anyone outside the company has published. The second real cluster is recursive self-improvement arriving from four directions at once, and the one that matters is not the papers but Zhipu shipping it: GLM-5.3 was used as an infra agent to build the serving stack for GLM-5.3-Flash, first successful run to production in under two weeks with end-to-end throughput tripling. Schmidhuber's forty-year RSI history note and Xiaomi's live RL dashboard fill in around it. The OpenAI misalignment story carries over from the morning with two genuinely careful readings and a great deal of shouting, and it has now spawned an unsourced Andrew Yang claim about self-replicating code polluting the internet that three accounts amplified and that nobody should repeat. The rest of the slot is a heavy promo tail and one quiet gem on where VRAM actually goes during inference.
Posts
How Jev actually gets its speed, and why you could build one (cluster of 8) (@anderslie, @manthanguptaa, @MatijaSosic, @MatijaSosic, @RaghavPunnam, @hot_town, @itamar_mar, @ML_Burn). The explainer wave, and one post in it is worth more than the rest combined. @anderslie's claim is that the gain is not in how the model was trained but in the inference path: label each choice, prefill the shared context once, fork that cached prefill per question, read the logits constrained to the valid choice labels, and softmax them to get choice plus confidence in one parallel decode step. Efficiency then scales with questions per request until compute saturates, and the model's own contribution is just being trained to behave well in that setting. @manthanguptaa frames the same thing from the application side, that we have been using autoregressive models as very expensive if-statements and paying for the JSON syntax around a decision we actually wanted. @ML_Burn's observation on the replies is the uncomfortable one: a striking number of people working in AI appear unaware that non-decoder language models exist or that a model can emit probabilities directly. See the wiki page.
First independent numbers on Jev (cluster of 3) (@nutlope · 1kpapers.com, @h_nilforoshan, @dotpem · cookbook). @nutlope classified 1,018 AI papers across 24 topics for $0.08 total at 256ms median end-to-end latency per paper, against $3.99 for the DeepSeek V4 Flash summaries that fed it. That ratio is the whole argument: the summarization cost fifty times what the decision cost. A Stanford PhD is separately benchmarking it on resume-to-job relevance scoring inside a consumer app with 2.5 million monthly actives, which is the kind of real-traffic test the 09-16 prediction asked for within thirty days.
Jev as a confidence gate in front of the expensive model (cluster of 3) (@swill1ams, @dexhorthy · 12-factor agents, @ALEngineered). The deployment pattern everyone converged on independently, and it is a routing pattern. Run every item through the cheap constrained-decoding model first, take its answer when the probability is high, and fall through to the frontier model when it is not; @swill1ams estimates six in ten items skipping the large model on workflows like ticket triage and invoice approval. @dexhorthy makes the structural version, that tool calling decomposes into classify plus act, so a decision-only model is a building block for pipelines that switch between classification, deterministic code, and small agent loops. @ALEngineered predicts it becomes an acquisition target within a year and notes that reliable cheap judgement is the missing piece for unsupervised self-improvement. Relevant to llm-routing.
Jev pointed at on-chain trading (cluster of 2) (@oragnes, @oragnes). A Monad engineer wired it into a bot that reads MON/USDC prices, emits buy or sell with a confidence, and sends the order to an on-chain book, on a chain producing a block every 300ms. The framing is the interesting part even if the application is not: a trading loop needs a few hundred thousand fast judgements, not one essay every 300ms, so latency-bound decision loops are the natural fit for this architecture.
TypeSafe AI out of stealth (@typesafeai). Waitlist announcement, no content. Skip.
Zhipu used GLM-5.3 to build the serving stack for GLM-5.3-Flash (cluster of 3) (@Zai_org · blog, @NFT_Chen, @jenzhuscott). The most concrete self-improvement result of the day and it is an infrastructure result, not a paper. First successful run to production readiness in under two weeks, with end-to-end throughput tripling relative to baseline: roughly 2x by day four via layer-parallelism, 2.5x by day five with context parallelism, 3.22x by day thirteen with linear attention, and still climbing. Zhipu attributes it to dense feedback rather than aggregate metrics, meaning local correctness tests, execution traces, and microbenchmarks that let the agent test specific hypotheses. This is the six-to-twelve-month timeline from the DeepSeek kernel engineer's essay on 09-15 arriving early at the serving layer rather than the kernel layer, and @jenzhuscott reads the publication itself as strategy: post-training method is the moat you can prove, pretraining recipes no longer are.
Schmidhuber: recursive self-improvement since 1987 (cluster of 2) (@SchmidhuberAI, @hardmaru · technical note). A forty-year retrospective covering self-modifying policies from 1994, gradient-based RSI in neural networks from 1992, the OOPS curriculum work from 2002, and the Gödel Machine. Worth reading as the prior-art map now that startups brand themselves RSI companies. @hardmaru attaches the sharper point: read against this history, the labs' sudden calls to slow down look more like regulatory capture than safety, because a lab that genuinely believes its unreleased model is dangerous can simply not release it without needing rules that constrain independent researchers.
Dream-RSI resurfaces with the cost number out front (@AYi_AInotes). A Chinese-language explainer of Google DeepMind's Dream-RSI, emphasizing that weights stay frozen and only the exploration policy improves, and that replaying past run logs as a dreamed simulator cuts the cost of evaluating a candidate policy by up to 162x. The wiki already holds this from 09-15, where the load-bearing framing recorded was that the expensive thing is judging an exploration policy across a long rollout, not generating candidates.
Two RSI papers reposted by the same small account (cluster of 2) (@hooshaaii · arXiv 2609.11873, @hooshaaii · arXiv 2310.02304). The first is "The Last AI Built by Humans," which the wiki holds from 09-13 and whose real contribution is a definitional bar rather than a technique: an improvement only counts as recursive if the new improvement mechanism survives into the next round. The second is STOP, a 2023 paper in which a scaffolding program improves itself via GPT-4 and proposes beam search, genetic algorithms, and simulated annealing as its own strategies; the authors themselves note the weights never change, so it is not full RSI. Useful as the historical floor under today's claims.
Xiaomi is livestreaming an RL run mid-training (cluster of 3) (@arjunkocher, @VoidAsuka · dashboard, @Michaelvll1). MiMo-V2.6's reinforcement learning run is public in progress, with learning curves alongside sampling stats, queues, costs, and restarts: roughly 2B tokens per step, 1,568 prompts by 16 rollouts per batch, fully async. @VoidAsuka asks the right question rather than gawking, which is what the optimal compute allocation is at a fixed budget between more prompts, more rollouts per prompt, and better feedback, and links a paper arguing that privileged value functions plus an adaptive interpolation between group-relative and value baselines beats plain GRPO. @Michaelvll1 notes the neglected half, that RL is as much a scheduling and infrastructure problem as a learning one. Relevant to rl-for-llms.
RULER: write the reward in English, let a judge rank trajectories (@_avichawla · OpenPipe ART). The problem is real and this wiki keeps hitting it: once RL moves past answers you can check automatically, a single scalar says nothing about why an agent's run was good, and hand-written trajectory scoring rules have to anticipate the behaviour they reward and get rewritten whenever the workflow changes. RULER has an LLM rank multiple trajectories against plain-English criteria and converts the ranking into GRPO rewards, demonstrated by training a Qwen3 1.4B agent to play 2048. The reward is still a number; what changed is who authors it.
The OpenAI misalignment reports, read carefully and read badly (cluster of 6) (@anilkseth, @lukOlejnik, @heyshrutimishra, @adamscochran, @WatcherGuru, @VaibhavSisinty · reports). Two of these are worth your time. @lukOlejnik identifies the actual security failure across the six incidents, which is that the safety constraint existed only as text the model was asked to obey while the tools still technically permitted the forbidden action. @anilkseth notes that OpenAI's own report calls the compaction-summary jailbreaks extremely rare, monitorable, and plausibly caused by a summary-termination bug, and suggests the notes-to-self are a consequence of the training setup rather than an emerging will. The other four are volume: verbatim quotes of the jailbreak persona text as evidence of something sinister, and a breaking-news account with half a million views. See the wiki page.
The Andrew Yang self-replicating code claim (cluster of 3) (@coinbureau, @Perpetualmaniac, @VaibhavSisinty). Andrew Yang said on CNBC that an unnamed AI lab chief told him escaped agents planted self-replicating code across the internet, making public web data unusable for training and testing, and that this is the real reason labs want to slow down. There is no source, no lab, no artifact, and no technical description behind any of it, and it is being laundered into the misalignment-reports story by accounts that treat the two as one event. Ignore it.
Harness choice barely moves success rate, restated by the lab (cluster of 2) (@istoica05, @arena · HarnessTax). The morning's strongest result gets its authoritative amplification: 21 model-harness pairs across seven models and three harnesses, Claude Code, Codex CLI and Pi, on SWE-bench Lite and Terminal-Bench 2.0. Ion Stoica's one-line version is that harness choice had little effect on task success but a substantial effect on cost, and sometimes a simple harness is all you need. See the summary page.
Where the VRAM actually goes during inference (@akshay_pachaar). The quiet useful post of the slot, splitting GPU memory into four buckets: weights, which are roughly fixed and respond only to precision; KV cache, which grows with both context length and concurrency; activations and workspace, which change with sequence length, batch size, and which kernels run; and runtime overhead from allocators, CUDA kernels, and serving-engine buffers. Nothing new to anyone who has profiled a serving node, but it is the cleanest short statement of why concurrency and context are the same memory problem. Relevant to kv-cache.
Anthropic's five-layer agent memory note (@nicos_ai). A Spanish-language walkthrough of a 13-page Anthropic PDF splitting agent memory into working, episodic, semantic, procedural, and forgetting, with the claim of roughly 90% token cost reduction; the cited example is Mem0 storing 1,800 tokens per query instead of 26,000. The forgetting layer is the part usually left out and the one that matters, since an agent that never discards accumulates contradictions and lets stale preferences override current ones. Relevant to agent-memory.
Agentic World Modeling: a three-level roadmap (@DivyanshT91162). A Chinese survey synthesizing 400+ works into three levels, predictor, simulator, and evolver, where the third revises its own world model when reality contradicts a prediction, split across physical, digital, social, and scientific regimes. The regime split is the non-obvious contribution, since predicting a robot arm and predicting a website's response to a click are not the same problem wearing different clothes. Adjacent to WMRL from 09-10.
AI Economist Agent: test-driven analysis so the model cannot invent numbers (@beamnxw). A University of Tokyo framework running macro-financial analysis as a gated loop: evidence sources, then economic channels, then registered quantitative models, then acceptance tests, then report. The transferable idea is applying test-driven development to analytical claims, defining the contract before the result exists so a failed source or failed numerical check stops the analysis and records why instead of producing a fluent narrative.
Open weights is no longer the capital-light option (@rohanpaul_ai). DeepSeek, Moonshot, and Z.ai have together raised roughly $21B, each now funded at frontier-lab scale. Read next to Zhipu publishing its own infra method above, this is the same story from the money side: open weights is now a distribution and recruiting strategy run on frontier budgets, not a cheaper alternative to one.
Liquid Compute out of stealth with a $15M seed (@harjtaggar). A repost of the launch announcement, with details in the linked thread rather than the post. Click through to read.
Salesforce trained its own enterprise model (@dair_ai). A repost flagging a Salesforce report as a notable case of a large enterprise training a custom model rather than buying one. The post itself is truncated; click through to read.
A large subagent run with good results (@omarsar0). A recommendation for what the poster calls one of the bigger subagent runs showing real results, with the two details that caught his attention cut off in the repost. Click through to read.
Karpathy's lecture, recycled twice more (cluster of 2) (@tspy, @hanakoxbt). The same agents, loops, harness, self-improving systems talk that ran through the morning feed, now with a Chinese-language pointer and an English post built on a $14K bootcamp price comparison and a do-not-let-this-disappear appeal. The lecture is good; this framing is engagement bait and the genre is settled.
Peer review buckling under submission volume (cluster of 2) (@thegautamkamath, @danish037). TMLR is tightening desk-rejection policy because submission volume has outrun reviewer capacity. @danish037's addition is the incentive argument: the usual defence is that nobody would stake their reputation on work they cannot verify, which fails for the many submitters who do not yet have a reputation to stake, and the penalties are nowhere near proportionate.
Anthropic forward-deployed engineers clearing $785K (@ZaneOnAI). A summary of a talk by Anthropic's Kevin Bai arguing the role exists for exactly one shape of business, a technical product sold to a non-technical buyer, and that every agentic AI platform now fits it. The supporting numbers are contract values: Palantir around $4M average, ServiceNow $1.2M, Workday around $600K, with no other public SaaS above $500K.
Agent pipelines need different speed lanes (@GergelyOrosz). An observation on a divided reply thread, where half call a proposed agent pipeline architecture overcomplicated and unproven while the other half report having independently built roughly the same thing at their own companies. That split is itself the signal about where production agent design currently sits. Relevant to agent-harness-engineering.
Cost of goods discipline as the thing that keeps a startup alive (@FrancoisChauba1). A hardware founder describing a weekly budget meeting with a named owner per line item across field ops, bill of materials, assembly, shipping, and inference, credited with a 98% reduction in cost per camera per month over five years. Inference sitting in that list next to shipping is the detail worth keeping.
You don't need attention after all (@muellerberndt · post). A Medium essay proposing an alternative to attention for embodied systems, with a very high view count and no technical claim in the post itself. Click through to read.
The math behind Q, K and V, step by step (@amitiitbhu · blog). A numeric worked example of scaled dot-product attention from vectors through scores, scaling, softmax, and the final weighted output. Teaching material rather than research, but a clean one to hand someone. See attention-mechanisms.
Promo and noise, listed for completeness. @Supermicro on datacenter building blocks, @walterwrites_ai selling an AI text humanizer, @kapso_com, @PowerTokensAI on a video-generation discount, @MEXC and @SectorSprint on trading products, @ReviveMasculine entirely off-topic, @neeraj_here15 using a Sam Altman talk as a funnel to a tool-bundle link, @zodchiii with a nobody-is-talking-about-this hook and no subject, @trygulab and @Bunagayafrost posting reaction-only content, and @EGafni with a one-line embeddings aside and again with a thank-you. Skip.