Summary
The morning slot is unusually thin on AI substance: 42 tweets captured, the curated retweet feed completely empty, and most of the volume from accounts posting political commentary rather than research. The strongest signal is a single linked essay from modem.dev on how coding agents actually navigate a repository, shared by an xAI engineer, backed by the unusual claim that the company's 680,000-line codebase is 99.9% LLM-generated. Second is a small cluster from Hugging Face's Elie Bakouch on open-model cyber capability, arguing that Kimi K3 may already sit at the top capability tier for offensive security while escaping the UK AI Safety Institute's evaluation because that eval caps at 100 million total tokens, a ceiling he considers far too low. Bakouch also pushes back on the idea of a coordinated global capabilities slowdown, calling it infeasible in practice and worth arguing against for that reason. A third thread is pure price competition: Grok 4.5's heavy subscription tier at $100 is being promoted as roughly half the cost of the Codex and Claude 20x plans with nearly double the rate limits, with Tim Sweeney offering a favorable comparison on programming-language-theory work. Beyond those three there is almost nothing: no papers, no benchmarks, no model releases, and a large volume of off-topic content from accounts that the AI keyword filter admitted on weak matches. A reader who only cares about research can skip this slot entirely and lose nothing.
Posts
How coding agents read your code, and how to write for them (@JonasBadalic, linking modem.dev). The strongest item in the slot. Ben Vinegar of Modem published part one of a series called Writing Code for Agents, and the credibility comes from the scale of the claim behind it: Modem's product is 360,000 lines of TypeScript application code plus 320,000 lines of test code, and 99.9% of it was generated by LLMs, starting back when Sonnet 3.7 was the leading coding model. The essay's core argument is that agents navigate a repository primarily by string search rather than by loading and understanding whole modules, which makes three things load-bearing that human-oriented style guides treat as taste: the names you choose, the types you define, and where you put your explanations. If an agent cannot find a symbol by grepping for the obvious term, it will not find it at all. Badalic's own framing is that this has been true all along and the contribution is quantifying the cost. The wiki's DSPy task-model separation page makes an adjacent argument at the interface level, that fixing a task contract makes the implementation swappable; this is the same instinct applied to source layout.
Kimi K3 may be at the top cyber capability tier and the standard eval cannot see it (cluster of 4, @eliebakouch plus follow-up). Bakouch summarizes a public exchange between a current and a former Anthropic researcher, and the technical detail worth extracting is an evaluation-design problem. He reports Kimi K3 might qualify as "Mythos tier" for offensive cyber capability, the label the discussion uses for the most capable class, but that this does not show up in the UK AI Safety Institute's evaluation because that eval is limited to 100 million total tokens, which he calls crazy low once input and cache hits are counted. The claim is that K3 is not reasoning-efficient enough to demonstrate its capability inside that budget, so a token-capped eval systematically underestimates models that need more thinking to reach the same result. Noah Lebovic, running an independent cyber eval, places K3 between Opus 4.8 and GPT-5.6 Sol. Lebovic's policy position, which Bakouch endorses, is that the model-access-for-defenders problem is not trending toward being solved by closed models, so open weights are preferable even at the top capability tier, conditional on those capabilities existing at all. Other threads in the exchange: Anthropic requires high spend before a model becomes cyber-capable, and current bad actors are still using Claude Code and Codex rather than anything exotic. This connects directly to today's digest thread on the 40-plus MCP CVEs and to the 07-26 breach reconstruction.
A global capabilities slowdown is not worth lobbying for because it is not feasible (@eliebakouch, responding to @tszzl). Roon said he would press a button to coordinate a global capabilities slowdown today if one existed. Bakouch agrees he would press it and argues nobody should lobby for it, on two grounds: it misrepresents how AI development actually works, and for some advocates the motivating factor is being able to say they tried if things go badly. He draws a line at the national level, saying a single-country slowdown is feasible precisely because it is enforceable, and that he would not press that one. The distinction is the interesting part, since it separates the desirability of a slowdown from its enforceability and notes they point in opposite directions.
Grok 4.5 pricing framed as the value leader (@ns123abc). The claim is that Grok 4.5's super heavy subscription runs $100, roughly half the price of the Codex and Claude 20x max plans, with nearly double the rate limits. Quoting Tim Sweeney, who says Grok 4.5 is very good for what he calls vibe math in programming language theory, keeping up with GPT-5.6 Sol, and more pleasant to use than Fable or Opus 5, which in his experience either generate superfluous documents in Claude Code or repeatedly prompt him to continue on the web. Worth reading as a preference report from one power user rather than an evaluation, but the rate-limit-per-dollar axis is a real competitive dimension and it is the one xAI is choosing to compete on.
Sam Altman wants a new kind of computer (@ns123abc, quoting @sama). Altman agreeing with an unnamed post that something "feels big" and saying he wants a new kind of computer. No product, no detail, no timeline. Noted only because it is the kind of remark that reads as a hardware-ambition tell in retrospect, and OpenAI's device work has been an open thread for over a year.
Benchmark readers and model users are different populations (@MillionInt). A one-line observation that people who look at benchmarks and people who actually use the models will never understand each other. Trite on its own, but it lands next to today's digest coverage of Opus 5 scoring 30.2% on ARC-AGI-3 against GPT-5.6 Sol's prior 7.8% record, a nearly fourfold jump on a benchmark explicitly designed to resist familiarity, which is precisely the case where the two populations would disagree about what the number means.
AI anxiety as a joke format (@stepango). An xAI engineer reacting to a video about a friend spending too much time talking to AI. No substance, listed for completeness of the AI-account feed.
Off-topic volume admitted by the keyword filter (cluster of roughly 30). The bulk of this slot's captured tweets are political and cultural commentary carrying no AI content: @brivael with eighteen posts spanning French politics, Musk commentary, and Fermi-paradox speculation, @spencerpratt with seven on Los Angeles municipal politics and water main breaks, @AustinJustice with three on license plate reader cameras, one @HouseGOP post about National Parents' Day, and three image-only posts from @heavypulp with no attachments captured. Skip.