Summary
The morning slot had no @bayesiansapien retweets, so all signal comes from the tracked AI-account feed, and it is dominated by two stories. First, Anthropic's "When AI builds itself" recursive-self-improvement report rippled across the feed: Anthropic's own Sholto Douglas described programming as already "IO-bandwidth limited" with models soon making only the hard calls, @ns123abc declared recursive self-improvement "happening," and Hugging Face's Elie Bakouch pushed back that the 8x-code-per-quarter plot may be confounded by which model version researchers actually had. Second, NVIDIA's Computex push continued at full volume: Cosmos 3 (an open omnimodal world model for physical AI), Nemotron 3 Ultra (550B/55B-active, 1M context, marketed for long-running agents), a worldwide AI-cloud buildout on Vera Rubin, and a fine-tuned Parakeet ASR hitting 97.7% Bahasa Indonesia accuracy. A coding-tools cluster ran underneath: Kilo put Nemotron 3 Ultra and StepFun's Step 3.7 Flash in free, brought back semantic codebase indexing, while Cursor shipped Canvas Design Mode and a context-usage report, and Bakouch charted Claude Code versus Codex feature parity (noting OpenAI's new "Dreaming" background memory). Standout non-cluster items: Prime Intellect joined the Nemotron Coalition with 2,500+ RL environments, Logan Graham tied a frontier-lab biorisk letter to Congress to his bedroom red-teaming origins, Google demoed passive smartphone-camera heart-rate monitoring, and Elon Musk pitched SpaceXAI satellites that run anyone's GPU or TPU. The political and personal noise from @brivael and @mlevchin is skipped.
Posts
Anthropic's recursive-self-improvement report, amplified across the feed (cluster of 3). Anthropic's "When AI builds itself" claims Claude now writes over 90% of Anthropic's code, engineers ship 8x more code per quarter than in 2021-2025, and a preview model improved on human researchers' next-step decisions 64% of the time, up from 22% in 2024 (@AnthropicAI article). Anthropic's own @_sholtodouglas framed the trajectory: "On my best days I literally feel IO-bandwidth limited managing the concurrent threads... very soon the models will only come to us for hard calls. Eventually they won't come to us at all." @ns123abc called it "Recursive Self-Improvement, IT'S HAPPENING." The skeptical note came from @eliebakouch, who warned the 4-week trailing-average plot hides the step effect of the new Mythos model and may be misleading if researchers were not yet using Opus 4.7 when measured. See wiki summary and today's digest.
NVIDIA's Computex blitz: Cosmos 3, Nemotron 3 Ultra, AI-cloud buildout, Parakeet ASR (cluster of 4). @nvidia introduced Cosmos 3, billed as the first open omni-model for physical AI, understanding and generating across text, image, video, sound, and robot action via a new mixture-of-transformers architecture, and topping open leaderboards for robot policies and scene understanding (blog). Separately NVIDIA pushed Nemotron 3 Ultra, a 550B/55B-active open model for long-running agents (5x faster inference, ~30% lower agentic cost, early adopters Perplexity, Palantir, ServiceNow), an expanding worldwide AI-cloud ecosystem on Vera Rubin, and a customer story where a fine-tuned Nemotron Parakeet ASR reached 97.7% Bahasa Indonesia accuracy (2.3% WER), cutting per-hour cost up to 90%.
Coding-tools cluster: free open models, semantic indexing, Canvas, Claude Code vs Codex (cluster of 4). @kilocode put two open models in Kilo for free: Nemotron 3 Ultra ("strongest US open-weights, 1M context, built for agentic work") and StepFun's Step 3.7 Flash (multimodal, 400 tok/sec, 256k context, near-Sonnet coding), and separately brought back codebase indexing, semantic search for "code you cannot name yet" (one query instead of four greps for the retry logic). @cursor_ai shipped Canvas Design Mode, letting you annotate UI elements directly to guide edits, plus an interactive context-usage report breaking down where tokens go across system prompt, tools, rules, and skills. @eliebakouch charted Claude Code versus OpenAI Codex feature parity (/goal, side conversations, subagents, the new "dreaming" memory), noting Codex produced a first UI in ~7 min where Claude Code took ~20 min, while flagging that "leading means shipping first, user experience is not only about this."
OpenAI "Dreaming" background memory for ChatGPT (@eliebakouch relaying @MTSlive). OpenAI launched Dreaming, a memory system that automatically synthesizes and updates user context in the background across conversations without explicit save requests, rolling out to US Plus and Pro users. It is the consumer mirror of the agent-memory research thread (recursive summarization, context compression) the wiki tracks under agent-memory.
Prime Intellect joins NVIDIA's Nemotron Coalition (@eliebakouch relaying @PrimeIntellect). Prime Intellect is contributing its 2,500+ open reinforcement-learning environments, the verifiers framework, Prime Sandbox, and NeMo Gym integration to the coalition (blog), arguing the hard part of frontier open models is no longer pretraining but post-training a base model into an agent that reliably completes tasks. This is the open-RL-environment substrate the self-evolving-agents papers in today's digest depend on.
Frontier-lab biorisk letter to Congress (@logangraham). Anthropic's Logan Graham tied a letter signed by Altman, Amodei, Hassabis and others (urging Congress to raise biosecurity) to his own history red-teaming LLMs for bioweapon risks since November 2022, arguing that what looked like overreaction in 2023 is now clear cause for caution, and that the real fix is deploying biodefenses. Feeds the responsible-ai thread.
Google passive heart-rate monitoring via smartphone camera (@GoogleResearch). A research system uses the front-facing camera to passively monitor heart rate during everyday smartphone use, claiming industry accuracy standards across all skin tones (blog). Health-sensing signal adjacent to the wearable/physiology foundation-model cluster topping Kurate this week.
Elon Musk pitches SpaceXAI satellites as hardware-agnostic orbital compute (@ns123abc). Musk said SpaceXAI satellites will let people run "whatever GPU or TPU they want," NVIDIA GPUs, Google TPUs, Amazon Trainium, or SpaceX's own future chips, plus its own AI software. Pure positioning for now, but it lands against SemiAnalysis's recent space-datacenter economics (4x terrestrial cost today, parity ~2040) covered in the 06-04 digest.
@brivael and @mlevchin: Skip. Political reposts, personal/cultural commentary, and a Lord-of-the-Rings disco remix. No AI research substance.