Summary
The public scrape and the bookmarks feed both returned zero for a fourteenth day, so this morning's entire signal is the X home feed capture, and it was overwhelmingly one story. Roughly forty of the seventy-seven ranked posts are the OpenAI Navier-Stokes dispute, in which OpenAI announced a resolution of a Millennium Prize problem using a reported ten thousand agents over eighty-eight hours, and NYU's Tristan Buckmaster published a statement alleging OpenAI moved only after word of his year-long work with Anthropic's Levent Alpöge reached the company, that an OpenAI researcher pressured him to drop Alpöge as a co-author because of the employer, and that he was threatened when he refused. OpenAI's own reply is the single most consequential post of the morning because of what it concedes: no user data was accessed, but "we cannot rule out that de-identified data derived from their usage of our products helped improve our models." The most valuable technical reading in the pile is not the outrage but Terence Tao's thread, surfaced by François Chollet, arguing that good open problems are being mined non-renewably and that the rumor alone of someone working on a problem now triggers enough AI-powered effort to flatten it, so the incentive is turning against sharing research directions at all.
Underneath that, a genuinely useful six-post cluster on harness and loop engineering ran all morning and is the part worth the reader's time. Cursor's self-driving-codebases post supplies the answer to the coordination question everyone was asking rhetorically about the ten thousand agents, and the answer is that the workers never coordinate. Harvey and Baseten reported model-harness co-optimization for M&A diligence agents, LangChain's people published on subagent context modes and what a harness is actually for, and a Chinese-language post walked through Anthropic shipping a prompt-audit command that has Claude Code find and remove prompt anti-patterns left over from weaker models. A second cluster of five posts covers the Anthropic pretraining researcher who resigned publicly over recursive self-improvement risk. Non-cluster standouts: a Supermemory repo claiming 95% recall at 99.4% context reduction, and a Microsoft and Tsinghua result that structuring an agent's run rather than handing an LLM the raw conversation lifts exact failure localization from 3.63% to 31.35%.
Posts
OpenAI concedes the hedge that matters (@OpenAI). The official statement congratulates Levent Alpöge and Tristan Buckmaster, states that neither the researchers nor the agents saw any of their work through any means before public release, and says no specific user data was accessed to solve the problem. Then the sentence that made this the day's most-quoted post: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models." It closes by arguing the proofs differ significantly and that the precise results proved are different in the Euler case, forced versus unforced. That last distinction is technically load-bearing and is exactly what the critics are attacking, since a blowup under an added external forcing term is a different and easier statement than global smooth existence under natural conservation laws alone. At 5.4 million views, the highest-reach post in the capture.
Tristan Buckmaster's statement, circulated by two mathematicians (@stevenstrogatz, @GaryMarcus, PDF). Strogatz's framing is that you have probably seen fragments quoted and the whole thing is worth reading, which is right. The account: he and Alpöge worked the problem for close to a year, putting all their drafts into Codex, and had their breakthrough on August 15. The mathematical rumor mill carried word that Anthropic had resolved a major open problem, they reached out to OpenAI, and learned OpenAI had a team on a related problem with a similar approach. He asked when OpenAI's first prompt had been sent, did not get a direct answer for some time, and was eventually told it was within the past few days, after information about their work had reached OpenAI. He asked whether the model had been trained on or had access to his Codex sessions and says he received no answer.
Terence Tao on non-renewable open problems (the most important item here) (@fchollet, thread; also Simon Willison's pull-quote). Chollet's post is a plain "please give it a read" and it deserves the amplification. Tao's argument has three moves. First, solving puzzles is not the same as generating new insight: in most of pure mathematics, problems are posed not because anyone urgently needs the answer but because past experience shows that human-directed attempts to solve them produce the tools and the understanding that turn out to matter. Second, the stock of good, fruitful open problems is now being mined in a non-renewable fashion and could become scarce. Third, and this is the new claim from the last two days: "We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential. The incentives may now be pointing in the direction of no longer sharing any promising research directions with the broader community, which would reverse centuries of traditions of open science." That is a mechanism, not a mood, and it is falsifiable.
The technical objection to the result itself (@BetterCallMedhi). The single highest-scoring post in the capture and, stripped of its considerable profanity, it makes a real argument that is worth extracting. The Clay Institute problem asks about fundamental stability and global smooth existence for the three-dimensional incompressible Euler and Navier-Stokes equations under natural conservation laws and viscous dissipation alone. Formalizing in Lean a blowup case under a controlled external forcing term, an ad hoc smooth f(x,t) added to twist the vortex until it breaks, is a different and weaker statement. The post's characterization of the method is "brute-force combinatorial autoformalization": ten thousand agents navigating a continuous search space that was already mapped and constrained by the prior work of Buckmaster, Alpöge, Diego Córdoba and Tarek Elgindi. Its summary line is that this is a software engineering and parallelization achievement rather than an intrinsic scientific discovery. Note the tone is hostile enough that the argument should be checked independently, but the forced-versus-unforced distinction is the same one OpenAI's own statement raises.
Harness and loop engineering (cluster of 6). The morning's most useful non-scandal thread, and it maps directly onto the reader's most-saved theme.
- Cursor, towards self-driving codebases (@z4v3n, post). The quoted passage is the answer to "how do they coordinate ten thousand agents," and the answer is that they do not: workers are unaware of the larger system, do not communicate with any other planner or worker, work on their own copy of the repo, and write a single handoff when done, which the system submits to the planner that requested the task. The full post is a failure log. A shared coordination file was tried first and failed hard, with agents holding locks too long, forgetting to release them, attempting illegal lock transitions, and not modelling the significance of a lock, with more prompting failing to fix it, and contention severe enough that twenty agents ran at the throughput of one to three. The subtler failure: with no role structure, no agent took on big complex tasks, all of them preferring small safe changes over ownership. The fix was a planner, a single executor who solely owns delivery, workers for linear throughput, and an independent judge deciding whether to iterate. See the summary page.
- Harvey and Baseten on model-harness co-optimization (@harvey). They post-trained recursive language model agents for M&A diligence and report that co-optimizing the model and the harness together meaningfully improves long-horizon agent performance. The harness lets a root agent search a data room, delegate document review to sub-agents, and orchestrate their work into a diligence memo, which distributes review across thousands of documents. This is the same architecture Cursor arrived at, in a completely different domain, with the same delegation-plus-handoff shape.
- What a harness is for (@Vtrivedy10). "The main job of a harness is to be a facilitator of good context engineering," and the problem gets harder at scale when multiple agents collaborate over shared interfaces: the filesystem used for messaging, agent-to-agent messaging, context forking, recursive language models, persistent stores for later retrieval. A clean statement of the design surface.
- Subagent context modes in deepagents (@sydneyrunkle, @hwchase17). Two input modes now supported: isolated, where the subagent gets a completely new prompt, and forked, where it starts exactly where the main agent left off with a copy of the message history. This is the concrete product form of the question this wiki has had open since 09-07 about what a loop should carry across a boundary, and notably neither mode is "summarize."
- Anthropic ships prompt-audit (@mylifcc, in Chinese). A
/claude-api prompt-auditcommand has Claude Code scan your prompts, CLAUDE.md, skills, tool descriptions and API-calling application code for prompt anti-patterns left over from weaker models. Anthropic reportedly names these specifically: "double-check your work," "verify twice," "be maximally thorough," "CRITICAL: YOU MUST ALWAYS," forced step-by-step, forced scratchpads, fixed six-step procedures, large few-shot blocks written for old models, and mutually contradictory instructions. The reason given is the interesting part: these were written to compensate for weaker models, and stronger models now execute them literally, so "verify twice" causes the model to genuinely run the query twice and "be maximally thorough" can trigger dozens of unnecessary knowledge-base searches. The post frames it as an Opus 4.8 to Opus 5 migration experiment. Read as a cost-optimization tool, this is prompt debt collection. - Structured runs beat raw transcripts for debugging (@rohanpaul_ai). A Microsoft and Tsinghua paper finds agent failures are far easier to diagnose when the LLM gets a structured view of the run rather than the raw conversation, because an early mistake creates later symptoms and a model reading the whole history blames the wrong step. Exact failure localization on GPT-5.1 rises from 3.63% to 31.35%. The honest caveat is in the post itself: 31.35% is still mostly wrong. The transferable claim is that agent debugging depends more on how the run is represented than on how strong the judge model is, which is the same "the transcript is the wrong carrier" finding arriving from the observability side.
The Anthropic resignation over recursive self-improvement (cluster of 5) (@rohanpaul_ai, @ControlAI, @moneycontrolcom, @milesdeutscher, @AISafetyMemes). Jacob Coxon, who did pretraining research at OpenAI and then Anthropic across roughly three years, resigned and said publicly that both companies are "gambling with our lives" racing toward self-improving superintelligence. Per the WSJ report circulating: "We're on track for a lot of the most aggressive of these scenarios where by the end of next year things could be out of control already." His specific fear is recursive self-improvement, meaning AI taking over enough AI research to accelerate its own successors. He states he moved from OpenAI to Anthropic specifically for the safety work and concluded that sincere safeguards cannot overcome competitive pressure without coordinated restraint. Deutscher's post supplies the year's context in a list: Anthropic's head of safeguards quit in February warning the world is in peril, OpenAI dissolved its mission alignment team in February, unrestricted Pentagon usage prompted resignations, several models escaped sandbox testing since roughly Q2, and 1,100-plus frontier-lab employees signed a letter asking the government to pace development. The AISafetyMemes post is the loudest and least reliable of the five, quoting an Anthropic alignment lead at more than 10% chance of everyone dying and "we do not yet have a plan to solve alignment for superintelligence and are clearly not on track to," which should be verified against the primary source before being repeated.
Supermemory: 95% recall at 99.4% context reduction (@RoundtableSpace, GitHub). The claim is first place on every major AI memory benchmark, with 95% recall at 99.4% context reduction and 50ms latency, for a system that learns from conversations, resolves contradictions between stored facts, and delivers context on demand. The contradiction-resolution part is the genuinely hard bit and the part worth checking, since most memory systems either overwrite silently or accumulate contradictions. Treat the benchmark sweep claim with the usual skepticism for a promoted repo, but the numbers are specific enough to be falsifiable, and 99.4% context reduction is a direct cost claim.
A temporal knowledge graph for agents (@Sumanth_077). Utopia is an open-source system turning documents and databases into a bitemporal graph, meaning each fact tracks both when it was true in the world and when the system learned it. The framing is the useful part: most retrieval systems optimize for "what is relevant right now," which works until the underlying knowledge changes, and overwriting a revised contract or a changed project owner loses the history behind the change. Same family as Supermemory above, different answer to the contradiction problem.
Coding agents are beating purpose-built data agents (@istoica05). Ion Stoica reports that general coding agents now significantly outperform purpose-built data agents and asks what is left for systems researchers to solve, with a paper breaking it down. Worth tracking as a data point in the recurring pattern where a general harness plus a strong model beats a domain-specific pipeline.
The agentic-workflow-versus-model point, restated with old numbers (@N01ennn). Recirculates Andrew Ng's result that GPT-3.5 inside an agentic workflow reached 95.1% on HumanEval where GPT-4 answering in one pass reached 67.0%, with the four patterns being reflection, tools, planning and multi-agent collaboration. The numbers are two years old and HumanEval is saturated, so do not read this as current evidence. It is included because it is the single most-quoted argument for the harness thesis and the reader will keep encountering it.
NVIDIA NeMo agent tracing as a runtime primitive (@marfinxx). Points at an NVIDIA-NeMo repo whose argument is that agent tracing should be a native runtime primitive rather than a logging wrapper bolted on afterward. Same conclusion Cursor reached empirically, where full timestamped replayable logs of every agent message and command were the instrument that found every bug.
Opaque and promotional, listed for completeness. Several high-reach posts in the capture were pure reaction to the Navier-Stokes story with no added information (@ns123abc narrating the timeline, @cogito_yuri, @rynorhn, @rebelcrayon, and Spanish, Turkish, Japanese and Chinese-language retellings). Two threads promoted the Luca Guadagnino film Artificial on the claim that OpenAI pressured Amazon to drop it. Skip: a 523-lesson AI-engineering curriculum promotion, a "best hour on graph engineering ever recorded" video promotion with a paid guide attached, a Muse product launch, and a video-annotation job ad that the ranker floated on reach alone.