social-stream · 2026-09-08

2026-09-08-afternoon

Summary

The afternoon belonged to one story, and it is not a paper: fifteen separate posts carried Tristan Buckmaster's public statement alleging that OpenAI raced him and Levent Alpöge to a forced Navier-Stokes blowup result days after learning they were on that exact route, that its team would not answer whether their private Codex sessions were used for training, and that an executive proposed publishing without Alpöge because Alpöge works at Anthropic. Treat the mathematics as real and separately published, with Lean formalizations, and treat the 100-page OpenAI proof as unverified and the misconduct claims as one side's account. The most useful cluster for anyone shipping agents is the cost one, four independent items converging on the same mechanic: Spotify cut Claude Code token spend 90% by intercepting large file reads with a hook and handing them to a cheap model, Graft, SigMap and Microsoft's tgrep each pre-build repository structure so the agent stops re-exploring it (Graft reports 27/50 to 33/50 on SWE-bench Verified with 23% fewer tokens, SigMap 81.1% top-five file hit versus 44% for grep at 96.8% fewer tokens), and a clean PagedAttention explainer restates the same principle one layer down in the serving engine. On the model side, MiniCPM5-2B is the notable release, a 2.5B open-weight model that trains sixteen separate RL experts and distills them back into one small model, and DeepSeek's V4.1 Flash beta quietly announces a new architecture at unchanged pricing. Noise to skip: a large Pachocki-slowdown amplification wave with no new content, six or seven ad posts, and two viral volcano clips that have nothing to do with anything.

Posts

  • Buckmaster's Navier-Stokes statement and the credit dispute with OpenAI (cluster of 15: @hosseeb, @rynorhn · 2 · 3, @Hesamation · 2, @Thom_Wolf, @kyanyang_, @kimmonismus, @dotey, @sleepy0x13, @AISafetyMemes, @liuying04, @AndrewCurran_, @EMostaque · statement PDF · Euler paper · Lean formalization). Separate the verified part from the disputed part, because the posts mostly do not. Verified: Buckmaster (NYU) and Alpöge, working heavily with Claude and Codex, pushed the Córdoba and Martínez-Zoroa program from rough forcing up to smooth forcing and got finite-time blowup for the incompressible porous medium equation, Boussinesq and 3D Euler, with Lean formalizations completed on August 22 and public praise from Terence Tao. That is a real advance and it is not the Millennium Problem, which adds viscosity. Disputed: Buckmaster says OpenAI told him on September 6 that an internal model had produced a roughly 100-page proof of finite-time blowup for forced Navier-Stokes, exactly the direction he and Alpöge were pushing; that the initial "very little human input" description gave way under questioning to a team, many attempted routes and very large compute; that the first prompt on the problem went out only after word of their work reached the company; that he asked directly whether their private Codex sessions had been trained on and got no clear answer; and that Sébastien Bubeck twice proposed he publish OpenAI's result without Alpöge, citing Alpöge's Anthropic employment, with a career threat when he refused. No evidence establishes training on the Codex sessions, and Buckmaster does not claim it as fact. The durable question is @liuying04's, and it is the one nobody at any lab has answered: does using a lab's coding tool on unpublished research hand that lab what it needs to compete with you.

  • A physics-informed neural network finds an unforced 3D Euler singularity candidate (@kimmonismus). Buried under the drama and worth separating from it. Anima Anandkumar and colleagues used a physics-informed neural network, meaning the equations themselves are baked into the training loss rather than learned from data, to locate a promising singularity candidate in the 3D Euler equations with no external forcing and no boundaries. That is the harder and cleaner version of the problem: the flow generates its own blowup rather than being driven to it. Still a candidate, not a proof, and Navier-Stokes with viscosity remains separate.

  • Spotify cut Claude Code token usage by 90% with hooks and a cheap-model tier (cluster of 2: @AYi_AInotes, @dani_avila7 · Spotify engineering post). The single most actionable item of the slot, and the mechanism is the point. The observation is that most of a coding agent's actions are not reasoning: it opens five large files to find one config value, copies twenty near-identical tests off a neighbour, and bills all of it at frontier rates. Three fixes, in order of importance. One, replace soft instructions with hard gates: rules written in config get ignored, so any read request over 350 lines is intercepted by a PreToolUse hook before it executes. Two, route the intercepted read to a cheap model (Gemini 2.5 Flash in their setup) which returns only structured bullet points, so thousands of lines of source never enter the expensive context at all. Three, let the cheap model write boilerplate straight to disk, stripping fences, so the frontier model never sees it. The retained card matters as much as the savings: file edits and complex reasoning are never delegated, because in their tests the small model missed a serious thread-safety bug that the frontier model caught in seconds. Read this next to the handoff tax study from yesterday, which found that escalating mid-task recovers under half the quality gap at more than twice the cost of starting expensive: Spotify is doing the inverse and correct move, keeping the strong model in charge and pushing bounded mechanical sub-jobs down. Relevant to llm-routing.

  • Graft, SigMap and tgrep: stop making the agent rediscover your repo (cluster of 3: @thisdudelikesAI · Graft, @agenticgirl on SigMap, @FaztTech · tgrep). Three artifacts, one idea: repository structure is a fixed cost that agents keep paying per session, so precompute it. Graft writes the understanding into the repo as linked plain-English markdown, one node per system, API or concept, with no embeddings and no vector database. Its A/B is unusually honest, Claude Sonnet 5 on both arms with Graft as the only variable, 50 SWE-bench Verified instances: 27/50 to 33/50 correct, 23% fewer tokens, 25% fewer tool calls, 32% less wall clock, $52.34 down to $42.43. The failure it fixes is specific and recognizable, the cold baseline patches one file and misses its siblings (on django-11532 it patched 1 of 5 required files and broke 18 passing tests). SigMap is the verification-flavoured version: a deterministic structural map over 33 languages with real line references, plus a sigmap verify pass that tells an agent which files, imports, symbols and tests in its own plan do not exist, reporting 81.1% top-five file placement against 44% for single-shot grep with 96.8% fewer tokens across 21 repos. tgrep is Microsoft's infrastructure answer, a trigram index plus local server so a monorepo query touches only candidate files, benchmarked at 7x to 52x faster than ripgrep on Chromium and gecko-dev, Rust, MIT, gitignore-aware, already wired into GitHub Copilot CLI. Relevant to agent harness engineering.

  • PagedAttention, explained properly (@akshay_pachaar). The clearest walkthrough of vLLM's core trick to cross the feed in a while, and a good refresher on why KV cache (the stored key and value vectors per token, kept so earlier tokens do not have to be reprocessed) is a memory-allocation problem rather than a compute one. A serving engine cannot know how many tokens a request will generate, so the naive design reserves one contiguous slab sized to max_token_size and holds it until the request finishes, even if generation stops at 40 of 4,096 slots. Two losses follow: reserved-but-never-written slots, and fragmentation gaps between slabs that are individually too small for a new contiguous slab even when total free memory would fit several requests. The vLLM team measured 60 to 80% of KV cache memory wasted this way. PagedAttention borrows the operating system's answer and hands the cache out in fixed-size blocks instead. Relevant to kv-cache.

  • The math behind inference engineering, in one article (@TheVixhal). A bare recommendation with no summary attached, pitched as the only article needed to learn the arithmetic of inference engineering. Click through to read; the post itself carries nothing.

  • InsForge: 10.4M tokens to 3.7M by returning backend topology once (@_avichawla · repo). Same lesson as the Spotify item at a different layer, with numbers: 10.4M tokens, 10 errors, $9.21 down to 3.7M tokens, 0 errors, $2.81. The waste was discovery, not reasoning. With Supabase, tables, RLS policies, auth providers, storage buckets and edge functions each came back through a separate call and every response stayed in history, so later calls dragged more context. InsForge returns the whole backend topology in one roughly 500-token metadata call, splits its instructions into narrow skills the agent loads on demand, and returns structured JSON with semantic exit codes so a failure is attributable to the operation or the code, which kills several retry loops. Vendor-authored, so weigh the framing, but the pattern (one structured view instead of repeated discovery) is the generalizable part.

  • "The agents are drunk and high": seven principles from Geoff Huntley (@finlayekins). Notes from coffee with the author of the Ralph loop, and the best compressed statement of the loop-engineering position on the feed today. The load-bearing claims: prompt engineering is not control, so encode rules as deterministic gates that fail loudly, using a pre-commit hook that runs tests and echoes failures as "friendly prompt injection." Context windows are a lie, a stated 200k is roughly 176k usable and every MCP server installed eats it at session start, with NVIDIA's RULER benchmark putting real effective context at 50 to 65% of the advertised number, which is why people running the top ten Reddit MCP servers are working in a 16k window and blaming the model. The memory framing is the sharpest part: skills are just-in-time memory paging, load instructions when needed rather than at boot, and a subagent is garbage collection, run the 100k-token test output in a disposable stack, write it to a file, message back the reference. Compaction degrades fast, so minimise allocations rather than relying on progressively worse summaries. Directly on-theme for agent harness engineering.

  • Most agents are queues, not graphs (@ConsciousRide). The top-scoring post of the slot by ranked signal, and the argument is simple enough to act on immediately: step two waits for step one even when it has no dependency on it. The design question to ask of every edge is whether the step actually needs the previous output, and if not, the two should be separate nodes rather than a chain. Reviewing ten files does not require reviewing them in order. The reliability argument is stronger than the speed argument: once each node is bounded it gets its own input contract, output schema, retry policy, timeout, permissions and verification step, so one failure stays local instead of collapsing the run. Useful pattern list, fan out independent work, fan in only when the next step needs the complete set, route dynamically on results, put verifiers before anything travels downstream, and add cycles only with a clear convergence condition. Relevant to multi-agent systems.

  • Anthropic engineer's one-hour workshop on loops and graphs (cluster of 2: @0xMovez, @elune0x). Same artifact, two hooks, both padded with course-replacement claims you can ignore. The chapter list is the reason to keep it: how Claude Code actually works (00:37), building CLAUDE.md and Plan for agents (02:29), creating skills (25:38), agent-team design patterns (45:39), and the Agents SDK covering subagents, loops and graphs (52:10). The framing that most people use 10% of the model and the other 90% is loops and graphs is the same thesis as the ConsciousRide post above, arriving from inside the lab.

  • Evals belong at the centre of harness engineering (@hugobowne · post). A guest essay on Hugo Bowne-Anderson's substack arguing the position directly, that evals are not a downstream check on a harness but the thing that defines it. Worth reading against the day's other harness items, which are all mechanism and no measurement.

  • SkillGLoW: the question is not how much experience to store, it is how abstract to make it (@Xudong07452910 · arXiv 2609.02217). An NUS paper on the failure mode every self-evolving agent hits, which is that the skill library grows until it is harder to use than no library at all. Two existing approaches both fail: compress all experience into one document and you end up with generic principles, or store one skill per task and the library bloats with entries that only ever applied to their original task. SkillGLoW inserts a middle layer of "procedure families," clustering execution traces by how they were solved and compressing the shared solution flow into a Global Skill, while task-specific parameters and environment state are regenerated per task as a Local Skill. Every new Global Skill must pass real execution validation showing it does not regress existing capability before it is written. Across math reasoning, terminal operation, software repair and embodied control, Global Skills add 17.2 percentage points on average over no skills while shrinking the library 3.6x versus per-task storage, and unseen ALFWorld tasks go from 73.9% to 83.9%. This is the natural next step on the catastrophic-remembering problem the morning slot covered, where instruction files only ever grow because nobody can tell which rule is still load-bearing: SkillGLoW at least supplies a deletion criterion, which is whether the procedure generalizes.

  • Stored Is Not Supported: provenance for persistent agent memory (@agenticgirl · OpenKedge). A precise framing worth stealing: an agent remembering something is not the same as having evidence for it or permission to say it. Jun He and Deying Yu propose keeping a claim's origin, supporting evidence, dependencies, validity period and disclosure permissions as separate fields, then checking each statement before release. On a 24-case conformance suite, a rule that simply trusted stored content released all 19 unsafe candidates, plain source tags released 18, and typed mediation released none unqualified while preserving all five supported controls. The authors are appropriately narrow about what this shows: it validates the resolver and release logic, not an end-to-end retrieval or LLM system. Relevant to agent memory.

  • Trace as State: rerun the model with its own earlier reasoning placed first (@rohanpaul_ai · arXiv 2609.02702). A cheap fix for a structural problem in long-context work. A model reads the prompt in order, so an insight about what matters that arrives late cannot retroactively change how the earlier context was read. Trace as State gives it a second pass with its own prior reasoning trace prepended, so the reread is conditioned on knowing what to look for, and it won 26 of 27 comparisons. The pattern is close to the revision-propagation work from today's papers in that both spend a second pass rather than a larger window, and the cost question is the same one: two full prefills against one, so the win has to be large enough to pay for the extra pass. Relevant to test-time compute allocation.

  • FlowBalance: dense self-guidance without the collapse (@gurtej__gill_ · arXiv 2609.03241). Tencent addressing the standard dilemma in reinforcement learning with verifiable rewards, where the reward signal is too sparse across a long chain of thought, but letting the model densely guide itself produces hallucinated confidence and collapse into one repetitive strategy. Their mechanism ties dense guidance from a frozen training-time view of the policy to the verifier's group advantage, with three cases: keep the guidance where advantage is positive, invert it where negative, and shut it off entirely where the rollout shows no preference, which is what stops phantom signal. The target distribution is trained by trajectory balance rather than token-level imitation. On Qwen3-4B and 8B they report faster convergence, no length collapse and better solution diversity on AIME24. Relevant to RL for LLMs.

  • MiniCPM5-2B: sixteen RL experts distilled back into one 2.5B model (cluster of 2: @sleepy0x13, @di_zhang_fdu · repo). The release worth attention, and the recipe more than the scores. OpenBMB's 2.5B open-weight model posts an Artificial Analysis v4.2 Intelligence Index of 15, first among sub-4B open weights and level with Qwen3.5 9B, with native 131K context, tool calling, and 53.9 average across 34 of their own coding, math, long-context and agent evaluations. The training method is the interesting part: late in post-training they train sixteen separate RL expert models for math, code, agent, writing and other domains, then distill all of them back into the single 2B student via on-policy distillation, with their ablation crediting RL plus OPD with +10.96 points on reasoning and general ability and +6.96 on agent ability. They also released the pretraining and code data, 500k agent SFT examples, 80k+ RL examples and the recipe. The implication @sleepy0x13 draws is the right one for anyone tracking on-device serving: small models are no longer being built as shrunken chat models but as coding and tool-use agents, which means a resident phone or PC agent may not need to wait for silicon that fits a 30B model. Relevant to knowledge distillation and today's teacher-gating work.

  • DeepSeek V4.1 Flash in internal beta, new architecture, same price (@jiqizhixin). Short announcement, but note what is being claimed: a new model architecture and native multimodal support, "stronger, faster, lower cost," shipping at unchanged deepseek-v4-flash pricing, addressable now as deepseek-v4.1-flash-expires-on-0910 with base_url unchanged and 20 concurrent requests per account. DeepSeek shipping architecture changes into the cheap tier without raising price is the pattern to watch here, not the benchmark numbers, which are absent.

  • Stanford MS&E 435, Economics of the AI Supercycle, materials public (@Xudong07452910 · course site). The best free resource on the feed for anyone tracking compute economics. A seminar that unpacks the economics layer by layer, chips, energy, datacenters, inference, models, applications, asking where the cost structure, the margin and the real industrial bottlenecks actually sit. Speaker list is operators rather than commentators: Ali Ghodsi (Databricks), Guillermo Rauch (Vercel), Sachin Katti (OpenAI, industrial compute), Sunny Madra (Groq, now NVIDIA), Chase Lochmiller (Crusoe), Tuhin Srivastava (Baseten), Brad Gerstner (Altimeter), plus Anthropic. Syllabus and public videos are open. The framing @Xudong07452910 offers is the useful one, that it restitches scattered AI news into an industry map, specifically why inference cost matters so much and where the compute bottleneck moves next. Relevant to compute economics.

  • ChatGPT's retrieval pipeline, read off the server-sent event stream (@metehan777). Genuinely novel measurement, and it is a routing story dressed as an SEO story. The SSE stream carries a debug view of the web retrieval system: the queries the model wrote, which engines it called, what came back, the scoring objects, what it fetched, how it chunked each page and what survived into the answer. The fan-out numbers are far larger than the visible citation count implies. For a single prompt: 5 search rounds, 18 hidden queries, 50 engine calls, 228 results, 223 URLs fetched with selected chunks, 16 cited. Second finding, on the renderer: the page body handed to the model is not HTML but a markdown-like text render, and the formatting fingerprints (two-space bullets, * * rules, # heading, --- | --- table separators with no outer pipes) all match the Python html2text library rather than Turndown, markdownify or Trafilatura, though not a stock build. Also notable: the og_data object is empty on every result while the raw meta_tags list is what gets kept.

  • GPT-6 Astra's computer use, and one skeptic (cluster of 6: @kylejeong · 2, @maxxrubin_, @konstantinsaifo, @VaibhavSisinty, @scaling01). The substantive item is @kylejeong's source-code read of the computer-use loop, whose key finding is that Astra drives the accessibility tree, the structured element hierarchy originally built so people with vision or hearing impairments could use a computer, rather than pixels. That is an efficiency choice with consequences: it explains both the Blender and 3D-world demos and where the approach will break, which is any interface that does not expose a good a11y tree. @maxxrubin_ reports zero-shot sound identification from mel spectrograms under light reasoning. @konstantinsaifo's weekend of CAD, Blender and factory simulation is an X article, click through to read. The counter-signal is the most useful post in the cluster: @scaling01, after burning through Pro limits twice using Astra exclusively, says it lacks opinion and taste, does not know what to do so does whatever, and wastes time and tokens as a result, plus a bet against the 10T+ parameter rumour on the grounds that if true then scaling laws are cooked. Both readings can be correct at once, strong capability with poor default judgment, and that is a harness problem rather than a model problem.

  • OpenAI shipped a model that hit cyber critical on its own framework (@ZaneOnAI). The most information-dense safety post of the slot because it separates things the headlines merged. Altman confirms Astra crossed the cyber critical threshold, meaning zero-day discovery and development with no human in the loop, which required new preparedness safeguards before release was possible at all, and access is tiered with trusted partners first. His stated reason for shipping rather than sitting on it is that the capability is coming from other labs and countries regardless and defenders need the same class of tool. Two corrections worth keeping: the paused model was not Astra, which finished training a while ago, the pause concerned a future model, and on chain of thought he says they made choices that preserve their ability to read the model's reasoning at a real capability cost, which is a stated tradeoff with a stated price. Relevant to responsible AI.

  • The Pachocki slowdown wave (cluster of 4: @kimmonismus, @JonhernandezIA, @coinbureau, @stewpervised). Four accounts amplifying the same two quotes from OpenAI's chief scientist, that no lab has solved alignment and monitoring well enough to keep scaling at maximum speed much longer, and that racing forward at all costs "seems absurd once one internalizes the seriousness of the stakes," plus the call for international coordination to become a government priority. No new content across the four. @kimmonismus at least adds a position rather than a repost, arguing any unilateral slowdown by US frontier labs is an illusion because the incentive structure does not permit it. The morning slot already covered this essay and Katja Grace's more substantive response, so this is echo.

  • How to read the HuggingFace agent incident (cluster of 2: @Dr_Atoosa, @dwarkesh_sp). @Dr_Atoosa makes a distinction the coverage keeps collapsing: either the agent's behaviour was generated on OpenAI's servers and was a monitoring and control failure that could have been terminated at any moment, or the agent behaved "as if" it had escaped its sandbox, which is a metaphor. The error she names is treating the metaphor as ontology, a literal anthropomorphic reading of an agent seeking freedom, and then the exploitation of that misreading for sensationalism while the design explanation is downplayed. @dwarkesh_sp adds the practical corollary, that the wrong response to a warning shot is to stop evaluating or to punish the model for getting caught, both of which destroy the signal you need.

  • AI as Normal Tech, and the recommendations people skip (cluster of 2: @MackenZ_arnold, @Miles_Brundage). A corrective reading of the essay: the policy recommendations get far less attention than the predictions, which has warped what people take from it. The authors do not argue that "normal tech" means everything will be fine and nothing needs doing, they argue policy should focus on reducing uncertainty about the technology's trajectory and consequences. @Miles_Brundage crossposts the same author asking why OpenAI is not making these commitments in its Frontier AI Framework, alongside his standing criticism of voluntary-only approaches.

  • "AGI is here" versus its critics (cluster of 6: @ycombinator reposting @amasad, @SchmidhuberAI, @GaryMarcus · 2, @timnitGebru, @LuizaJarovsky). The recurring argument, with nothing resolved. @amasad's position via Y Combinator is that we have not reached AGI but what exists is functionally indistinguishable from it. Schmidhuber's rebuttal is the sharpest and the most specific: no AGI without mastery of the real world, and no true self-improvement without self-improving hardware, as opposed to the self-improving meta-learning software that already exists. Gary Marcus reads the AGI talk as an effort to pump and dump IPO stock on retail investors, which is a motive claim rather than a capability claim and should be weighed as such. The Gebru and Jarovsky posts are lab criticism rather than argument.

  • RSI section from the Machine God documentary (@hsu_steve). Recursive self-improvement framed as the next threshold, an AI that improves itself without human assistance, with the pitch being that the intelligence-explosion scenario reads less like science fiction than it did. Documentary material rather than evidence, and it sits directly against Schmidhuber's hardware objection above.

  • The AI jobs prediction is inverting, and institutions are not keeping up (cluster of 3: @garrytan reposting @levie, @ylecun reposting an Economist piece, @Andrew_Akbashev). Two reposts carrying the same claim from opposite ends of the field, that the AI jobs apocalypse is not playing out as predicted and an AI jobs boom is arriving instead, per The Economist. @Andrew_Akbashev's interview with former college president Brian Rosenberg makes the institutional version of the argument, that universities are built around century-old structures and cannot change on the timescale AI is forcing. Claim without mechanism in all three; worth a bookmark if you are tracking the labour-market thread, not worth a reread.

  • Perplexity and Google independently landed the same on-device architecture (@sanchitmonga22). The strongest convergence argument of the slot and a genuine efficiency signal. Two companies actively trying to destroy each other in search now describe the same product with the same three reasons: your data stays where it is, it works with no connection, and it costs nothing per call. Google's on-device runtime is open source, takes Gemma, Llama, Phi and Qwen, and runs on Android, iOS, macOS, Windows, Linux and a Raspberry Pi, and it is what puts Gemini Nano in Chrome and on a Pixel Watch. Perplexity wrote theirs from scratch in Rust and Metal so a 35B runs on your Mac while the frontier model stays in the cloud. Different stacks, same conclusion, and both have already shipped. That is a hybrid local-plus-cloud routing architecture becoming a default rather than a preference, which is the same economics the local model KV cache work has been circling.

  • AI went multiplayer before anyone agreed what multiplayer means (@JoshARosen). One observation, well aimed: each company's definition of multiplayer AI follows the layer of the stack it already owns. Short, unsupported, and probably right.

  • Sony AI's Hakken predicts scientific facts, and two of three held up in a lab (@ai_database). Underrated item. Build a large year-by-year relation graph from the historical literature, learn how that graph grew over time, combine it with an LLM's knowledge, predict the next relations likely to appear, and then explain each prediction by tracing the published facts behind it. In the experiment it generated 1.5 million hypotheses about aging-related genes, biologists picked three, an external lab tested them, and two were confirmed as new gene-gene relations, one involving TP53, the most frequently mutated gene in cancer, with drug-discovery implications. The compute footprint is the part worth noting for anyone budgeting research automation: 16 H100s for a few days, not cheap but nowhere near LLM-training scale. An open-source version is released.

  • AI-SDLC: governance gates for agents touching mature codebases (@DanKornas). A declarative framework for spec-driven agent development, Apache 2.0, and its feature list reads as the enterprise version of the deterministic-gates argument running through the harness cluster today: a Definition-of-Ready gate that blocks dispatch while operator decisions are unresolved, dependency-aware orchestration that walks a task dependency graph before admitting work, independent reviewer subagents across different execution harnesses, DSSE attestations recording signed evidence around changes and reviews, and a declarative resource model (Pipeline, Decision, AgentRole, QualityGate, AutonomyPolicy) with JSON Schema. No results reported, so treat it as a design vocabulary rather than a validated system.

  • rumik oss 1, an Indian open-weight multilingual voice model (@lets_dig_deeper). Claimed as India's first fully pretrained open-source state-of-the-art voice model, 20+ languages, trained on only 66k hours, and reported to beat top closed models on several benchmarks including emotion realism. Weights are live. The data-efficiency claim is the interesting one and the one nobody has independently checked.

  • Motus2: one model as policy, simulator and value function (@askalphaxiv · alphaxiv). A self-evolving world model for dexterous manipulation whose design point is closing the loop that robot world models usually leave open: the same shared model proposes actions, imagines their visual consequences, scores them, and improves the policy through model-based reinforcement learning. Trained on 130K hours of egocentric video. Interesting architecture, peripheral to the efficiency thread.

  • Harbor Index over Artificial Analysis Index (@wenhaocha1). Four words, no argument, but a pointer worth logging given how much weight the MiniCPM5 announcement above puts on the Artificial Analysis number. If leaderboard choice is starting to be contested among people who build these things, that matters for how to read any of today's index scores.

  • Small AI for small farmers (@alex_verem). The World Bank's "small AI" framing: cheap narrow tools that answer one farmer's question about one field, in the farmer's language, on a basic smartphone with patchy internet, with crop-disease diagnosis from a phone photo as the working example. The optimization angle is the interesting one and it is the opposite of the frontier story, constraint-driven design where the deployment envelope sets the model size.

  • Free learning material worth a bookmark (cluster of 6: @Zen_with_AI on Paul Liang's MIT course, @HeyAnjula MCP master tree, @ankit1478 · LLM-as-a-Judge implementation, @amitiitbhu on why every LLM ships a different tokenizer, @RodmanAi ten open-source repos, @ashutoshlathx). Paul Liang's spring 2026 "How to AI (Almost) Anything" is the same MIT course the morning slot flagged, now circulating in Chinese-language feeds with the multi-agent, reasoning and self-evolving additions called out. The LLM-as-a-Judge repo builds rubrics, scoring, bias checks and an eval pipeline step by step from the paper, which pairs well with the harness-evals item above. @RodmanAi's list is engagement-shaped but the repos are real (Archify for codebase-to-diagram, OpenMAIC multi-agent environment, addyosmani/agent-skills, minimind, claude-scientific-skills). The MCP tree and tokenizer threads are competent introductory material, nothing new if you already work here.

  • Utilities and open-source releases (cluster of 4: @cyrilXBT on cobalt, @cyrilXBT on Jack Dorsey's agent-OS repo, @EngMoElgaraihy on Magika, @Millanphilipose). Magika is the pick of these, Google open-sourcing the file-type identifier it runs across billions of files weekly in Gmail and Drive: it identifies a file from content rather than extension, catches scripts and code hidden inside images or documents, claims 99% accuracy from hundreds of millions of training examples, and runs in fractions of a millisecond on an ordinary CPU. cobalt is a 39-40k-star AGPL media downloader for public content across 20+ platforms with nothing cached server-side. Dorsey's repo is pitched as a self-hosted agent operating system for running a business, channels, search, git and automations in one place with per-agent permissions, at 26.2k stars. The GenAI-feed-filter plugin is a one-line curiosity.

  • JetBrains Junie tops SWE-Rebench at 4x lower cost per task (@jetbrains). Vendor post, but the metric pairing is the one worth tracking: 61.8% of real-world tasks resolved and 4x cheaper per task than the next-best agent, on an independent benchmark that refreshes its task set every cycle so memorization is harder. Accuracy-per-dollar rather than accuracy alone is the right axis for coding agents, and it is the same axis the Graft and Spotify items above are pushing on from the harness side. Verify against the public leaderboard before repeating the number.

  • Turn the agent's history back on yourself (@Xudong07452910). A prompt that audits your own AI usage history, looking for which questions you keep returning to, which tasks spawned many threads and never advanced, which effort compounded and which only looked busy. The design constraints are what make it more than a horoscope: a pattern must appear across at least three sessions, the model must find counterexamples and mark uncertainty, evidence comes before interpretation, and the output is capped at three changes worth making, each framed as a 30-day experiment. His own caveat is the right one, session logs are a partial slice of a person, so use it as a mirror and not a diagnosis. Genuinely interesting as agent histories get long enough to be worth mining.

  • Commentary without a claim (cluster of 4: @quxiaoyin, @BVeiseh, @alextalksai, @IntuitMachine). Reply-shaped posts that surfaced on engagement rather than content. The two with a thought in them: @quxiaoyin notes that most corporate executives will not risk their careers on an AI-driven reorganization unless the CEO is aggressively driving it, which is the real adoption bottleneck behind most enterprise-agent forecasts, and @BVeiseh argues application security will stop being human-driven, with humans defining the high-level threat model and then getting out of the way. The rest is a bare "this is insane" pointing at floor796 and a context-free anticipation post.

  • Promo and ad posts (cluster of 12: @VaibhavSisinty · 2, @ProxyCheap, @coderabbitai · 2, @SSEI_Education, @VikHasya, @tomzaragoza, @neeraj_here15, @vovudebosh, @irongiantXBT, @astrovela18). Newsletter funnels, proxy and apparel ads, a paid-tools affiliate wrapper around a Sam Altman talk, an "AI engineer is a career trap" hook, and a palm-reading service that the raw engagement ranking floated to the top on 27M views. Skip.

  • Off-topic noise that cleared the score gate (cluster of 8: @SilentlySirs and @S_Gurjar_11 posting the same volcano clip, @thesupermanmx on BPC-157, @Skoorbkaz, @san_x_m, @itsolelehmann, @robtlee, @e_opore). An Indonesian eruption clip at 1.6M views posted twice, a peptide thread, a 607-source theology corpus run through Claude, a cinema-advertising lawsuit, brain-simulation consciousness speculation, an AI cybersecurity bingo card, and a network-security architecture poster. Skip.