Summary
The morning slot has no curated reposts, and the AI handle feed is dominated by one story: the fallout and resolution of Anthropic's invisible Fable 5 throttling. A cluster of three accounts (@ClaudeDevs, @eliebakouch, @ns123abc) covers the walkback from inside, from the research community, and from the breaking-news angle respectively, all confirming that flagged frontier-AI-development requests will now visibly fall back to Opus 4.8 and return a refusal reason on the API. The second real thread is cheap-and-fast inference shipping as product: Google's DiffusionGemma (surfaced via @tinygrad quoting Sundar Pichai) claims 4x faster text generation by being compute-bound, and @kilocode launched Coding Plans built on MiniMax M3, pitched as near-Opus-4.8 quality at a tenth of the price, against the backdrop of Roo Code shutting down and Copilot moving to usage-based billing. Standout non-cluster items: @cursor_ai's Bugbot is now 3x faster and 22% cheaper, Anthropic's @_sholtodouglas amplified Poetic's $50M raise for high-accuracy long-horizon enterprise automation, @GoogleResearch posted a new framework for auditing machine unlearning and differential privacy, and @nvidia ran a GTC Taipei news cycle (Vera CPUs at NYSE, a 248-Hopper Stanford SuperPOD). A batch of pure political and personal posts from @brivael and others is noise.
Posts
Anthropic makes Fable 5 safeguards visible (cluster of 3) (@ClaudeDevs, @eliebakouch, @ns123abc). @ClaudeDevs announced that starting this week, requests flagged as targeting frontier LLM development will visibly fall back to Opus 4.8, the same mechanism as the cyber and bio safeguards, and the API will return a reason for any refusal. The framing is an admission: "we went with invisible safeguards... and that was the wrong tradeoff." @eliebakouch (Hugging Face) welcomed the reversal, saying his biggest concern was hiding the degradation from users and the paranoia that would create, and used it to argue why open models and open research are critical so good-faith researchers always have access to the best AI. @ns123abc covered it as breaking news, summarizing that researchers found invisible guardrails "secretly degrading users' AI research," called it "secret sabotage," and that Anthropic apologized. This is the social-feed face of today's lead Industry Pulse and Global View items in the daily digest.
DiffusionGemma: 4x faster text generation, open weights (@tinygrad quoting @sundarpichai, blog.google). Google released DiffusionGemma under Apache 2.0. The article explains the core idea: a normal autoregressive LLM is memory-bound when serving a single user (the time to generate one token is the same for 1 user as for 256, because the cost is loading weights from memory), so the individual user gets no latency benefit from batching. DiffusionGemma flips this by spending that idle compute on one user, starting from a 256-token random "canvas" and refining all tokens in parallel over several passes, the way image diffusion denoises an image. It is compute-bound instead of memory-bound, so a single user's generation scales with added compute. @tinygrad's own comment was a hype check: Google is the largest compute owner in the world, and "AI is not a race, it's a decentralized revolution that will take decades."
Kilo launches Coding Plans on MiniMax M3 (@kilocode, blog.kilo.ai). Kilo now lets users buy external coding-plan subscriptions directly with their existing Kilo credit balance, starting with MiniMax M3, which it claims benches near Claude Opus 4.8 at a tenth of the price. The post frames it as insurance against vendor churn, citing two recent shocks: Roo Code shut down and took its workflows with it, and GitHub Copilot moved every plan to usage-based token billing on June 1, keeping the flat fee but removing the unlimited ceiling. The pitch is that engineers now run two to five coding tools at once and each is a moving pricing target.
Cursor's Bugbot is 3x faster, 22% cheaper, finds 10% more bugs (@cursor_ai, cursor.com). The code-review agent's biggest update yet: 90% of runs now finish under three minutes. A new pre-push
/reviewmode lets you run Bugbot and Security Review locally before opening a PR, and it deduplicates against the GitHub or GitLab review so the same diff is not scanned twice.Poetic raises $50M at $500M for hallucination-free enterprise automation (@_sholtodouglas amplifying @markiewagner). Anthropic's Sholto Douglas boosted Markie Wagner's launch of Poetic, described as a system that executes complex multi-hour tasks with 99%+ accuracy and 10x fewer tokens than agents, raising $50M at $500M from Kleiner Perkins, Founders Fund, First Harmonic, and Genius Ventures. The pitch is that code is too brittle and agents are too unpredictable for work like anti-money-laundering and fraud investigations, the high-stakes back-office processes that run the global economy.
Google Research: auditing machine unlearning and differential privacy (@GoogleResearch). A new framework, Regularized f-Divergence Kernel Tests, for auditing machine unlearning (verifying a model actually forgot data it was asked to delete) and differential privacy. Google says it is more sensitive to localized data shifts and needs fewer samples than traditional tools. Relevant to the responsible-AI privacy-audit thread.
NVIDIA GTC Taipei news cycle (@nvidia). Several items: Stanford deployed "Marlowe," a DGX SuperPOD with 248 Hopper GPUs, opening large-scale compute to 500+ researchers across all seven schools; NVIDIA Vera CPUs are powering NYSE market infrastructure with Redpanda and HPE; and the NVIDIA AI Podcast hosted Mistral CTO Timothée Lacroix on Mistral's open-model philosophy and its Forge customization framework under the Nemotron Coalition. Mostly product and partnership signal, no research substance.
OpenAI weighing "drastic" price cuts and Stargate questions (@ns123abc). Two posts: OpenAI is reportedly considering drastic price cuts to win the user war with Anthropic, with Altman saying costs "have become a huge issue" ahead of the IPO, and a claim that the $500B Stargate 2.0 Ohio figure is a full-buildout estimate rather than a commitment, with the actual near-term plan around 800 MW by 2028. Treat as rumor-grade until corroborated, but consistent with the inference price-war theme in today's digest.
Opaque X long-form repost (@magicsilicon). Pushkar Ranade linked an x.com/i/article on the parallels between jetliners and transistors (the largest and smallest things we build), but the article body could not be fetched (expired X cookies). Click through to read.
Personal and political posts (@brivael, @bcherny, @lynnmartin). Skip. @brivael's feed is French-language philosophy and political commentary with no AI research substance, @bcherny posted a greeting from Code with Claude Tokyo, and @lynnmartin posted about the Knicks.