Summary
The morning belongs to a launch, and the interesting part is that the launch is about deleting text rather than generating it. Diogo Almeida, who co-authored InstructGPT and the RLHF work at OpenAI, came out of two years of stealth with Jev, a model that makes decisions over a predefined option set and never writes a sentence, priced at $0.042 per million input tokens with output tokens free. It is a cluster of ten posts across four languages and most of them read like coordinated promotion, but two do not: a developer who ran roughly 5,000 requests for about $2 across classification and model routing, and a skeptic who reduced the whole thing to a really smart switch statement. Underneath the noise, the morning's genuinely useful technical content is a cluster of three on inference internals: a CUDA worklog that derives online softmax from first principles, a clear explanation of why a running model changes character halfway through a request when prefill ends and decode begins, and an X article breaking down HBM system architecture down to the DRAM operating principle. The agent-memory conversation had its densest day in weeks, with four separate posts, and the sharpest line came not from any of the launches but from a researcher noting that human testers get faster with practice while agents slow down as their memory notes grow. Two long-horizon agent systems landed, Microsoft and Shanghai Jiao Tong's Argus with 1,548 wall-clock hours and one human intervention every 40.7 hours, and Google's WikiSkill, which compiles an agent's execution history into validated reusable skills. Salesforce's Koa paper circulated with its abstract screenshotted: an enterprise agent model built by post-training open-weight Nemotron-3-Super-120B with GRPO on tasks expanded from the company's own declarative agent config files. And the day's most honest post came from a theory researcher who spent four weeks and over ten billion tokens pointing agents at the deletion channel, an open problem from the 1960s, did not resolve it, and said the experience was emotionally taxing enough that he is stepping away from agent-assisted theory work.
Posts
Jev launches, a model that only makes decisions and never generates text (cluster of 10: @CompleteSkeptic, @rohanpaul_ai, @aakashgupta, @k_grajeda, @Alisina_ai, @VaibhavSisinty, @shiri_shh, @trikcode, @MaxForAI, @namcios). The founder is Diogo Almeida, co-author on InstructGPT and the RLHF work at OpenAI, previously at Google Brain, out of two years in stealth. The pitch: every time an LLM makes a decision inside software, it produces that decision by writing a sentence or a JSON object one token at a time and your code parses it back out. Jev deletes that layer. Give it a question plus a predefined set of options and it returns structured decisions and probabilities directly, computed in parallel rather than token by token. Claimed at 20-200x faster and 40-400x cheaper than the LLM route, at $0.042 per million input tokens with output tokens priced at zero, which follows naturally from a model that emits almost no output tokens. The training method is called RLCD and no paper accompanies the name. The model explicitly cannot write code or generate natural language. The framing "hallucinations are impossible because it cannot speak" is a definitional claim, not an accuracy claim, and several of the amplifying posts present it as the latter. Ten accounts, four languages, near-identical talking points, all in one day, which is what a promotion campaign looks like. Covered at length in today's routing summary.
The one independent Jev data point, and it is worth more than the other nine posts combined (@MichaelLee04). A developer with early access reports roughly 5,000 requests for about $2, spanning classification, model routing, intent detection and steering. His framing is the useful one: Jev is a decision-making primitive that sits between deterministic code, which is too dumb, and LLM calls, which are too slow and expensive, and he expects to make several Jev calls around every LLM call in his product. He also brings a real prior: he had been using DeepSeek Flash and more recently the GPT-5.6 Luna family as low-latency classifiers for exactly this work, on questions like whether a conversation output contradicted something the user said earlier, or whether enough silence has passed to send a proactive message. That is a genuine before-and-after comparison from someone who was already paying for the thing being replaced, which is the only evidence in the whole cluster that is not vendor-supplied.
A two-hour open-source answer to the two-year stealth launch (@harshagundal). Released as Qwen-2.5-1B-RLCD on HuggingFace, claimed at 5x faster on-device inference for type-safe JSON workloads, demoed on an M4 MacBook. The argument underneath is the interesting bit and it partially deflates the launch: every LLM can already batch-infer every key of a JSON simultaneously and emit probabilities over a set of candidate categories, with no new training required. If that is right, the decision-only form factor is a serving-time restructuring available on existing open weights rather than a new class of model, and the proprietary version's advantage is the training method and nothing else. Nobody has benchmarked the two against each other.
A CUDA worklog that derives online softmax instead of asserting it (@athletic_coder, worklog). Anshuman Mishra takes softmax from the most naive kernel to a warp-per-row implementation, quantifying cost at every step. The derivation that matters: stable softmax subtracts the row maximum before exponentiating, which appears to require a full pass to find that maximum first. It does not. If you are partway through a row with running maximum m1 and running denominator L, and you meet a larger value m2, multiplying the whole running sum by e^(m1 minus m2) corrects every already-stored term at once, because e^a times e^b equals e^(a+b). Then add 1 for the new maximum's own term. One multiplication repairs the entire history, which is what lets the max pass and the denominator pass be fused. The same identity reappears across space: with one warp per row, each of 32 lanes holds a local maximum and denominator, and each partial is rescaled by e^(local max minus row max) before the sum reduction. That is precisely the mechanism that makes tiled attention possible, since you cannot process attention in blocks unless you can merge partial softmax statistics from different blocks. Global memory traffic falls from 16MN to 12MN bytes, and the naive arithmetic intensity is computed explicitly at about 0.25 FLOPs per byte, which tells you it is a bandwidth problem before you write any code. See the summary.
The clearest short explanation of why prefill and decode are different machines (@agenticgirl). During prefill the GPU receives a large block of prompt tokens at once, so most of the work becomes efficient matrix multiplication with good weight reuse and enough parallelism to push toward compute-bound. Then decode starts and each active sequence advances by exactly one token, so the model revisits its entire weight set and an ever-growing KV cache (the stored key and value tensors that save recomputing attention over past tokens) to produce a tiny amount of new output. At modest batch sizes memory traffic dominates arithmetic, which means a GPU showing low FLOP utilization during decode may be doing exactly what the workload allows. The post then derives most of the inference-engineering canon from that one transition: batching amortizes weight movement, grouped-query and multi-query attention shrink the KV state decode must carry, PagedAttention handles the allocation mess created by variable-length caches, FlashAttention attacks attention data movement, chunked prefill exists because a large prompt otherwise interferes with ongoing decodes, and prefill/decode disaggregation lets the two phases use different scheduling, batching, parallelism and sometimes different GPU pools. It lines up exactly with the disaggregated serving chapter ingested yesterday, which argued colocating the two phases forces one tensor-parallel degree on two workloads that want opposite ones.
A full system architecture breakdown of HBM, from DRAM first principles to vendor differentiation (@siliconcodesign). An X long-form article covering the operating principle of DRAM, the difference between DRAM and logic processes, base die versus core die features, and where SK Hynix, Samsung and Micron actually differ. The article body could not be fetched, so this is the pointer rather than the content, but the topic sits directly under the wiki's memory hierarchy page and specifically under the capacity-versus-bandwidth argument that ran through the 4-hi HBM analysis earlier this week. Worth clicking.
Argus runs for 1,548 hours and asks for help once every 40.7 hours (@jiqizhixin, paper, code). Microsoft and Shanghai Jiao Tong University, open-sourced. The reported numbers are 27 research campaigns, 1,548 wall-clock hours, one human intervention every 40.7 hours on average, and a work duty cycle of 95.1% to 98.7%, including solving a 20-year-old math problem. Four design principles: evidence-driven, where every claim must be backed by verifiable evidence; self-evolution, where agents learn from their own research history; multi-agent collaboration with separate roles holding independent contexts but a shared workspace, specifically to avoid local hill-climbing; and core-vertical decoupling, where the core handles permissions, evidence submission and human boundaries while each vertical defines what counts as valid evidence in its domain. The attached Figure 1 shows the runtime as a Manager with authority over tasks, verticals and stage transitions, driving a Planner, Engineer and Reviewer over a shared workspace holding a knowledge wiki, an event and decision log, and code artifacts, with a LifeSupervisor, backlog, budget, daemon and memory persisting across bounded missions. The surrounding evidence cards report SWE-Bench Pro 78 against a direct baseline of 59 (1.41x), AARRI-Bench 76.8 against a paper baseline of 68.3, nanochat B200 and H100 wall-clock times just under human reference, SOL-ExecBench rank 6 with seven top-3 finishes across 101 kernels, and a math-data score of 28.0 against Arbor's 20.83, Claude's 8.33 and Codex's 6.25. The authors label the card row explicitly as breadth evidence rather than a single normalized leaderboard, which is unusually honest labelling.
Google's WikiSkill compiles an agent's execution history into validated reusable skills (@beamnxw). The described loop: run tasks, preserve raw traces, consolidate recurring failures and successful strategies into a wiki, propose one atomic skill update, validate it, keep or roll back. The asymmetry is the design choice worth noting: skills can be rolled back and the wiki never is. Successful strategies, recurring failures, rejected edits and skill impact history all survive into the next iteration, so the accumulated record is append-only while the executable artifact derived from it is revertible. That is the separation between evidence and policy that the wiki's harness literature has been circling, stated as an implementation rule. The poster's framing, an experience compiler for agents, is the right one.
The sharpest sentence about agent memory this week, and it is a negative result (@ManlingLi_). Quoting a definition of continual learning as "the process by which an apprentice develops expertise on the job," she flags the finding that caught her: human testers got faster with practice, while agents generally slowed down as their memory notes grew. Her reading is that accumulating skills can simply burden an agent with its own memory, and that multi-scale abstraction is the missing piece, meaning experience has to be compressed to the right level for each task rather than merely retained. Read alongside today's continual-learning paper, which pushed retention across 100 sequential tasks from 1.2% to 34.9%, the pair says something uncomfortable: accumulate into weights and it degrades, accumulate into notes and it slows you down, and nobody currently has a third option.
MEMENTO, an agent memory layer open-sourced by an Anthropic engineer (@0xWast3). MIT licensed, described as 41 memory types, 9 write gates and 6 decay curves. The design rule is that nothing is remembered simply because it happened: every fact must earn a slot and every slot carries an expiry. The write path is intake, where a claim is split from its source so the fact and who said it never fuse; a gate stage running nine checks including whether the fact already exists in different wording; a conflict rule; scope separation, so facts about a person, a project and a preference go to separate stores; and decay, where anything unread long enough loses weight until retrieval stops surfacing it. The conflict rule is the part worth stealing. A new fact that contradicts an old one does not overwrite it. Both stay, timestamped. The poster's argument is that every memory system on the market silently overwrites, so yours has been quietly wrong for months and you cannot tell because the old version is gone. Keeping the trail lets you find the exact turn a fact changed and what caused it. The 41-types and 9-gates framing is likely inflated for the post, but the non-destructive conflict handling is a real design position and the agent memory page has consistently found deletion and contradiction to be the unsolved half.
Anthropic's agent-memory playbook circulates again with its token-cost claims (@adiix_official). The same five-layer stack that circulated on 09-15, now with a claimed 90% token-cost reduction attached to the headline. The layers: working memory (the context window, where most agents stop), episodic memory (timestamped interaction logs, so the agent recalls that the deploy failed Tuesday because of a typo in the migration script), semantic memory (facts and relationships as a knowledge graph, where "user prefers TypeScript" lives and does not expire with the session), procedural memory (an approach that worked becomes a reusable skill), and forgetting (an agent that never forgets accumulates contradictions and keeps recommending restaurants in the city the user left). The quoted figures are secondhand from the PDF: Mem0 storing 1,800 tokens per query instead of 26,000, and Snowflake adding one ontology layer for 20% better accuracy and 39% fewer tool calls. Treat the numbers as unverified. The 90% headline in particular is the poster's, not the document's, as far as the captured text shows.
Salesforce Koa, with the abstract screenshotted (@omarsar0, paper). The attached image is the paper's first page and it is more precise than the tweet. Koa is built by post-training the open-weight Nemotron-3-Super-120B foundation model with reinforcement learning using GRPO (Group Relative Policy Optimization, an RL method that compares a group of sampled responses against each other instead of training a separate value model). It is trained on public and synthetically generated data with no customer data. The distinctive component is a simulation-to-reward pipeline that expands workflow specifications into persona-conditioned multi-turn tasks, with rewards grounded in successful tool use for data-dependent requests. For enterprise domains those specifications are written in Agent Script, Salesforce's declarative language for building Agentforce agents; for public tool-use domains the workflow structure is synthesized directly. Reported to improve over its open-weight base across public tool-use, agentic-reasoning and CRM benchmarks with the clearest gains on multi-turn tool use, surpassing a strong proprietary baseline while remaining below the strongest frontier models, which the abstract states plainly. The thesis in one line: specification-driven RL is a practical path to specializing open-weight foundation models for enterprise agentic work. The config files you already wrote to define your agents are the training data.
HuggingPapers surfaces today's continual-learning paper with the headline number (@HuggingPapers). The attached card is the paper's own figure and it is clearer than the tweet: three anchors feeding one weight update, labelled data anchor (replay), function anchor (matching the previous distribution) and weight anchor (constraining toward previous weights), plus the low-rank allocation choice shown as shared LoRA, where the update is applied to the weights and the next adapter reuses the same structure, against merged LoRA, where the update is merged into the base weights and the next adapter is freshly initialized. Composing three anchors with merged LoRA lifts average final retention after 100 tasks from 1.2% to 34.9%, a 28-fold improvement. The dataset is on HuggingFace. Ingested today as the composition paper.
Google Research introduces Retrieve-for-Train, replacing autoregressive inference with a diffusion model for search slates (@GoogleResearch, blog). The framing is a direct cost claim: complex AI search currently runs heavy autoregressive inference to produce a slate of expert-level results, and swapping that for a lightweight diffusion model gives instant slates at a fraction of the cost. Only the announcement text was captured, so the mechanism is unverified here, but the shape is the day's recurring one. This is the same move as Jev, made by a different lab in a different domain: take a decision that was being produced by generating tokens sequentially and produce it in one parallel shot instead. Two independent instances on one morning is worth flagging even though neither has an independent benchmark attached.
Four weeks, ten billion tokens, an open problem from the 1960s, and an honest negative result (@DimitrisPapail). Dimitris Papailiopoulos pointed AI agents at finding the capacity of the deletion channel, a coding-theory problem open since the 1960s. They did not resolve it. He describes it as four weeks of GPU hammers and more than ten billion tokens combining old and new approaches. In a follow-up he notes this was also the longest he has had agents collaborate on one problem, three weeks, and that communicating agents seem to add a dimension of capability rather than simply more tokens, which is a claim worth tracking because it is the first one from someone with no incentive to make it. In a third post he says he is stepping away from agent-assisted theory work for a while because the experience has been emotionally taxing. Read all three together. It is the most credible account of the actual texture of agent-assisted research anyone posted today, precisely because it reports a failure.
SOAR, a teacher that invents problems it cannot itself solve (@CrazyShyyt). The framing is overheated but the mechanism described is coherent. When a model is trained on problems with a 0% initial success rate there is no reward signal, so it learns nothing, a failure the poster calls the edge of learnability. The usual fix is to supply millions of easier human-written examples as a bridge. SOAR splits the model into a Teacher and a Student: the Teacher looks at the impossible problems and generates new easier ones as stepping stones, feeds them to the Student, and is rewarded if the Student improves on the impossible test. The claim that would matter if it survives scrutiny is that the Teacher does not know how to solve the stepping stones it creates, and that the structure of the practice problems matters more than the correctness of their answers. No paper link was attached, so treat it as a description rather than a result.
Perplexity claims a hundred million dollars a year from replacing DynamoDB with two engineers and agents (@AravSrinivas). The stated build: an in-house key-value database for fast web content fetches, built by two engineers plus hundreds of persistent agents over two months, with migration expected to save up to $100M annually. No technical detail, and the saving is a projection rather than a realized number. It is recorded because it is the second concrete claim this week that a small team plus persistent agents can attempt an infrastructure rewrite that previously needed a department, and because the saving is framed entirely as a cloud-vendor bill rather than a headcount one.
"The End of Prompt Engineering," amplified by someone with a system to sell (@the0xbt). The described argument is that instructing agents is finished as a skill, because a prompt is a wish and a wish repeated three hundred times is still a wish, whereas boundaries written into a versioned file are read by every agent before every action. At swarm scale, the claim goes, three hundred agents cannot share a vibe but they can share a file, so the job shifts from writing instructions to being a boundary architect who defines what never happens and audits what gets asked. No paper link is attached and the post's second half is a pitch for the poster's own AGENTS.md setup with three hundred Kimi K3 readers and "zero escapes in 41 days," which is an unverifiable self-report. The underlying position, that declarative constraints compose where imperative prompts do not, is one the wiki's harness literature supports, but it is being asserted here rather than shown.
Karpathy's graphs thesis, via a bookmark (@0xnicc0). Saved to bookmarks today, summarizing an hour-long Karpathy talk as a progression from LLMs to prompts to agents to graphs, with the graph as the endgame and everything before it a stepping stone. The stated substance: the point is not tweaking text in a chat window, it is building the state machine that isolates errors, preserves context, and runs parallel agent workflows that never reset. The post is heavily engagement-framed ("people pay thousands for bootcamps that teach half of this," "don't let this vanish from your feed"), which is a reason to read the talk rather than the tweet. Treated properly in today's Media Zone.
A KV cache engineering explainer, saved and worth its own read (@_avichawla). An X long-form article on KV cache engineering for LLM serving: why the cache grows, the twelve ways models and serving engines reduce it, what each technique actually saves, and the trade-offs that decide which one fits. Saved to bookmarks today. The article body could not be fetched, so this is the pointer, but the topic is the reader's core area and a systematic twelve-technique taxonomy is exactly the shape the KV cache page organizes around. Covered in today's Media Zone.
A 277-page free textbook that covers the chapters most books skip (@KirkDBorne, book). Foundations of Large Language Models by Tong Xiao and Jingbo Zhu, covering the lifecycle from pre-training data and architecture through alignment and, unusually, inference acceleration and system-level scaling. Those last two are the reason it is listed rather than skipped. It also circulated on 09-15 from a different account, so this is a second independent surfacing.
Schmidhuber argues current LLMs are not creative because nobody implemented his 2008 theory (@SchmidhuberAI, paper). The referenced work is "Driven by Compression Progress," which proposes that subjective beauty, novelty, surprise, interestingness, curiosity and creativity all reduce to one principle: an observer finds data interesting in proportion to how much its compressor improves by processing it. The claim that LLMs have not implemented it is fair on its face, since no frontier training objective rewards compression progress as an intrinsic signal. The post is the day's highest-ranked item by reach-normalized engagement, which says more about the account than the argument.
Skip. An agent-evals course launch, a six-month roadmap thread for becoming an evals engineer, a Sam Altman Stanford talk repost circulating from two accounts with identical copy, a video-editor hiring ad, three trading-bot promotions in the long tail, a Sora-diffusion-transformer explainer thread with no new content, and the day's Anthropic-IPO speculation. No research substance.