Summary
The morning slot's own scrape was empty again: no curated reposts and no tracked-handle tweets in the 24-hour lookback. The morning home-feed capture carried the signal, covering the US afternoon and evening of 09-29, which was OpenAI DevDay. The biggest cluster is DevDay itself (cluster of 7): Dots agents, the Ultrafast speed tier, GPT-6.1 Sol in the Arena, and a CNBC stumble where OpenAI's CFO called Dots "Muse." The most useful technical signal for cost work is a cluster of 3 on decisions and routing: SGLang shipped a native /v1/decisions endpoint that turns any served model into a classifier, a TypeSafe skill teaches Claude Code to build on Jev, and Anthropic's cost-optimization skill ranks token-saving levers. A hardware item stands out: General Compute is splitting requests across chips, with GPUs doing prefill and Cerebras doing decode. Safety and governance form a third cluster (the White House accord, Altman on pacing, Gary Marcus on PBS). Nvidia's open coding-agent stack (Nemotron 3 Ultra, OpenShell) and its Physis-Lang world-model captions round out the standouts.
Posts
OpenAI DevDay (cluster of 7). (1) Sam Altman launched Dots, always-on agents that run on their own cloud computers and work in the background (@sama · post). (2) OpenAI's Ultrafast tier: up to 8x faster generation, about 300 tokens per second in Codex, up to 6x in the API (@OpenAI); The Decoder reports it costs six times as much. (3) Dots roll out to Pro, Business Premium and Enterprise, against Meta's Muse, which is free to start but US-only (@rohanpaul_ai). (4) CFO Sarah Friar called Dots "Muse" on CNBC (@StockSavvyShay). (5) GPT-6.1 is live in Arena's Agent Mode, where models run long web, filesystem and terminal tasks (@arena). (6) A reported DevDay detail: OpenAI researchers let models continuously optimize the computer-use harness itself (@TheTuringPost). (7) On more open-weight gpt-oss models, Altman said "if we can do something useful there, we will" (@omarsar0); one poster read it as OpenAI trying to capture open-source inference because "GPUs are the only moat" (@not_ellington). Full breakdown in the DevDay page.
Decision endpoints and cost levers (cluster of 3). (1) SGLang turned Qwen3.8-27B into a multimodal decision model that beat Pokemon FireRed's Elite Four with sub-100 ms decisions. The new native
/v1/decisionsendpoint makes any LLM or VLM a classification and scoring model, and/v1/systemonemakes Jev-like open models work with TypeSafe's SDK (@sgl_project). Jev is TypeSafe's decision model that returns calibrated probabilities instead of text. (2) TypeSafe released an official skill (about 1,300 words) that tells Claude Code to read the live docs and keep rules, maths and lookups in normal code, using Jev only for judgment calls (@alex_prompter). (3) Anthropic's/claude-api cost-optimizeskill profiles where tokens go, ranks levers by savings, applies free wins first (prompt caching, input hygiene, batching) before tradeoffs (effort, model), and measures one diff per lever against your eval (@dani_avila7 · skill). See the Raschka and decision-API page.General Compute: prefill on GPUs, decode on Cerebras (@rohanpaul_ai). A neocloud for non-Nvidia chips makes its first move: a large Cerebras purchase funded by $400M of debt. Each request is split. GPUs read the prompt (prefill, compute-heavy). Cerebras writes the answer (decode, memory-bandwidth-heavy), fast because the weights sit in on-chip SRAM rather than HBM. Details in the hardware page.
Nvidia's open coding-agent stack (@suraj_sharma14). Nemotron 3 Ultra, a 550B open model with 55B active parameters for long-running agents, ships with open data and 173B tokens of GitHub code. OpenShell 0.1.0 is an open sandbox that policy-locks files, network and credentials. A new benchmark, SWE-Serve, found one in three patches that passed local tests failed in live serving. The post's hype framing ("Nvidia just killed Claude Code") is noise, but the artifacts are real. Nvidia also named OpenShell as the governance option for OpenClaw Enterprise (@NVIDIAAI).
Governance and safety (cluster of 4). (1) Six labs signed the White House accord: internal controls, external audits and board review, all voluntary (@rohanpaul_ai). (2) Altman on CNBC: "We are pacing our progress, which includes sometimes not training the model" (@rohanpaul_ai). (3) Gary Marcus on PBS: self-regulation is not enough (@GaryMarcus). (4) OpenAI seeks at least $30B at about $1.4T pre-money; Altman reportedly ruled out a 2026 listing over safety concerns (@rohanpaul_ai).
Physis-Lang: physics reasoning in video captions (@NVIDIAAI · project). Nvidia researchers generate captions that explain why and how a scene unfolds physically, then use them to fine-tune video world models and to enrich prompts. With Cosmos 3 Super and Nano it ranks first and second on the Physics-IQ Verified image-to-video leaderboard.
Distributed training on AMD (@PyTorch). A PyTorch Conference talk on GPU-initiated networking inside TorchTitan, from AMD's RCCL work, aimed at communication overhead. vLLM's conference presence includes Simon Mo's keynote (@PyTorch).
Agent infrastructure launches (cluster of 4). Replicas V3 (cloud agents for Claude Code and Codex), Agent37 (sandboxes claiming 94% lower cost), Browser Use Ultrafast (flight comparison for $0.004 in under 20 seconds), and an agent-payments startup raising $35M (via @harjtaggar, @ycombinator, @garrytan, @shmidtqq). Cost claims are vendor claims.
Learning resources (@techNmak). A list of visual explainers for transformers and inference, led by Transformer Explainer and Brendan Bycroft's LLM Visualization. Useful for teaching, nothing new.
Skip. Jev-plus-Sonnet design-hack threads and harness "10 repos" lists are engagement posts. A viral "AI launches nukes" study summary cites 2025-era models. A Dot demo failure became a stock-picking post.
Today's full digest: cere-bro 2026-09-30.