Summary
OpenAI DevDay dominated the day, and all of it arrived through the morning home-feed capture: Dots always-on agents, the Ultrafast speed tier at about six times the price, GPT-6.1 in Arena's Agent Mode, and a CFO slip on CNBC. The afternoon and evening slots were empty, which fits the US overnight clock, so there is no cross-slot cluster today. The most useful signal for cost work is the decisions-and-routing cluster: SGLang's native /v1/decisions endpoint turns any served model into a classifier, and Anthropic's cost-optimization skill ranks token-saving levers by savings. The single-slot standout is General Compute splitting each request across chips, with GPUs doing prefill and Cerebras doing decode from on-chip SRAM. Nvidia's open coding-agent stack has real artifacts under a hype wrapper, and the governance cluster (White House accord, Altman on pacing) is mostly talk. For the full picture, read the daily digest and the Media Zone.
Posts
- OpenAI DevDay (cluster of 7) (@sama · Dots) [morning]. Dots agents run in the background on their own cloud computers. The Ultrafast tier reaches about 300 tokens per second in Codex, and GPT-6.1 went live in Arena's Agent Mode (@OpenAI, @arena, @rohanpaul_ai, @TheTuringPost, @omarsar0, @StockSavvyShay). More in the DevDay page.
- Decision endpoints and cost levers (cluster of 3) (@sgl_project · cost skill) [morning]. SGLang's
/v1/decisionsreturns calibrated scores from any LLM or VLM, and TypeSafe's Claude Code skill leaves rules to code and uses Jev only for judgment calls (@alex_prompter, @dani_avila7). See the decision-API page. - General Compute prefill/decode chip split (@rohanpaul_ai) [morning]. GPUs handle the compute-heavy prefill and Cerebras handles bandwidth-heavy decode, funded by $400M of debt. See the hardware page.
- Nvidia open coding-agent stack (@suraj_sharma14 · @NVIDIAAI) [morning]. It bundles Nemotron 3 Ultra (550B total, 55B active), the OpenShell sandbox, and SWE-Serve, which found that one in three locally passing patches failed when served live. Ignore the "killed Claude Code" framing.
- Governance and safety (cluster of 4) (@rohanpaul_ai · @GaryMarcus) [morning]. Six labs signed a voluntary White House accord, and Altman said OpenAI sometimes chooses not to train a model. OpenAI is also reportedly raising at least $30B at about $1.4T pre-money (@rohanpaul_ai).
- Physis-Lang physics captions (@NVIDIAAI · project) [morning]. Captions that explain the physics of a scene are used to fine-tune video world models. With them, Cosmos 3 takes the top two spots on Physics-IQ Verified.
- Distributed training on AMD (@PyTorch) [morning]. A PyTorch Conference talk showed GPU-initiated networking in TorchTitan, built on RCCL, to cut communication overhead. vLLM's Simon Mo also gave a keynote (@PyTorch).
- Agent infrastructure launches (cluster of 4) (@harjtaggar, @ycombinator, @garrytan, @shmidtqq) [morning]. The launches were Replicas V3, Agent37 sandboxes, Browser Use Ultrafast, and a $35M agent-payments raise. The cost figures are vendor claims.
- Visual explainers list (@techNmak) [morning]. It lists Transformer Explainer and Bycroft's LLM Visualization. Good for teaching, nothing new.
- Engagement threads [morning]. Skip. These include Jev design-hack threads, "10 repos" lists, a stale "AI launches nukes" summary, and a Dot demo failure turned into a stock pick.
- Empty slots [afternoon + evening]. No reposts, tracked-account tweets, or linked content, as expected in the US overnight.