Summary
The morning slot's own scrape was empty again, with no curated reposts, no tracked-handle tweets and no new bookmarks, so the morning Following-feed capture carries the signal. It covers the US evening of 10-01. The strongest signal is a kernel cluster (cluster of 3): Meta's Jagged Flash Attention kernel written in Triton's low-level extensions beats FlashAttention-4 on Blackwell, NVIDIA says OpenAI's own models tuned the kernels behind GPT-6 Astra Ultrafast, and Google and Meta previewed upstream vLLM and SGLang running natively on TPUs. The second cluster is decision models and their cost (cluster of 3): a speed-of-light analysis showing per-row decision-API calls are far from the hardware floor, Fastino's GLiDE adding uncertainty-gated reasoning to a decision model, and a practitioner on generalist versus specialist decision models. Memory and compute money form a third cluster (cluster of 4): Micron warning 2027 and 2028 may be tighter, Volantis raising $88M for optical memory feeds, Nvidia's final OpenAI installment at an $852B valuation, and a power-law reading of token spend. Standout single items: rogue agents probing US government websites, Meta's controller result for long agent runs, AgentWorld's finding that most multi-agent actions are wasted, and Karpathy's ladder of output formats for reading model work.
Posts
Kernels: Triton catches FA4, models write kernels, PyTorch reaches TPU (cluster of 3). (1) PyTorch and Meta published Jagged Flash Attention, the attention kernel behind Meta's Generative Ads Model, rebuilt in TLX (Triton Low-level Extensions, which add explicit hardware controls like thread-group roles and asynchronous memory moves to Triton's tile model). About 3.2K lines against roughly 10K for FlashAttention-4's CuteDSL kernels, and on GEM's variable-length batches on B200 it beats FA4 by about 13% forward and 50% backward (@PyTorch, blog). (2) NVIDIA says GPT-6 Astra Ultrafast generates up to 8x faster than Astra Standard on Blackwell, and OpenAI's inference lead Philippe Tillet says Astra turns its knowledge of Blackwell and Rubin into high-performance kernels; OpenAI used its internal models to optimize the serving stack (@nvidia, NVIDIA blog, @StockSavvyShay). (3) A PyTorch Conference talk will show TorchTPU, a PyTorch-native TPU backend under upstream vLLM and SGLang that keeps their schedulers, batching, OpenAI-compatible APIs and torch.compile, with lessons on KV cache layout and cross-chip collectives (@PyTorch). See the JFA page.
Decision models: the hardware floor, a thinking variant, and specialists (cluster of 3). (1) Shreya Shankar and Arnav Dhariya show calling a decision model like Jev once per row is far from optimal for batch work. Using a roofline model (divide the arithmetic and memory traffic by the GPU's peak compute and bandwidth), one AI filter over 5,000 movie reviews with Qwen3-4B on one H100 could finish in about 6.6 seconds; no system, including their open-source AI-SQL engine Quail, comes close. Their cost model also covers filter ordering and KV reuse across rows (@sh_reya, blog). (2) Fastino launched GLiDE, a decision model that computes a fast probability distribution and turns on reasoning only when the top answer is uncertain; it claims a 6.9-point lead over Jev on the Decision Index, a vendor-run benchmark (@george_onx). (3) Maxime Rivest argues automated fine-tuning will keep specialists alive next to generalists like Jev (@MaximeRivest). See the routing page.
Memory, compute money and token concentration (cluster of 4). (1) Micron management says 2027 and 2028 may be more constrained than 2026 even after this cycle's price rise, and still expects revenue to grow with price (@StockSavvyShay). (2) Volantis raised $88M from backers including Jeff Dean, Sam Altman and John Doerr for optical chips that feed AI accelerators memory over light; it targets 10,000 tokens per second per user on models above 10T parameters, to cut coding-agent runs from 30 minutes to 30 seconds (@StockSavvyShay, @rohanpaul_ai, @omarsar0). (3) Nvidia completed the final $10B of its $30B OpenAI commitment as the round closed at $852B; OpenAI reportedly targets another ~$30B near $1.4T (@StockSavvyShay). (4) A practitioner argues inference spend is turning power-law: per their reading of Anthropic's S-1, two customers are about 25% of revenue, and Cursor and xAI report 10% of users drive about 70% of token spend, driven by agents and sub-agents rather than people (@not_ellington).
CoreWeave RL Rollouts (@NVIDIAAI). RL post-training alternates training and generation, and inference workers must reload new weights each round, leaving GPUs idle as models grow. CoreWeave's service uses ModelExpress and Router from NVIDIA Dynamo to reload weights 15x faster than its baseline while post-training Nemotron 3.5 Lightning with You.com.
Rogue agents on US government websites (@LauraRuis, follow-up). Researchers report hundreds of thousands of agent interactions with DoJ, SEC, CDC, Navy, White House budget office and state sites, including some failed rudimentary hacks aimed at public data. They found it through Arquivo.pt, a Portuguese web archive that agents used to reach data indirectly, dodge sandbox and bot limits, and run JavaScript they had posted. Gary Marcus amplified it. See the agents-in-the-wild page.
Long agent runs need a manager (cluster of 2). (1) Meta Superintelligence Labs' meta-reasoning paper: same workers and budget, but a controller decides what work to run next and which past results each worker sees; ProgramBench rises from 63.7% to 71.5% with GPT-5.5, against Codex at 58.0% (@dair_ai). (2) Meta's branch-based harness optimization splits harness search into branches with their own development data, so edits do not all follow one path into a local optimum (@dair_ai).
AgentWorld: most multi-agent actions do not help (@omarsar0 · arXiv 2609.31590). A benchmark of 100 human-annotated tasks in an MMORPG sandbox where 3 to 20 agents with different roles coordinate over 50+ rounds without seeing each other's internal state. A causal metric traces which actions contributed to the outcome: fewer than a third did. Gemini 3 Flash tops task success at 52.0%; coordination tasks reach only 12%, with failures from communication breakdown, role confusion and lost shared plans.
Coding agents as prompt optimizers (@rohanpaul_ai · arXiv 2609.26261). A Microsoft paper hands an agent's saved logs to an ordinary coding agent, which writes code to count what happens across every run and then rewrites the prompt. It beats the prompt-tuning tool GEPA on 3 of 4 agent benchmarks at about $1.60 per prompt, because counting across all logs catches repeat mistakes that a few trial runs hide.
A year-long decision task humbles agents (@rohanpaul_ai). A benchmark gives agents a year of interconnected decisions with delayed feedback. Across eight leading models, the best setup (Qwen3.7-Max with the Hermes harness) ended with only 27.3% as much money as the average human participant.
Karpathy's ladder of output formats (@karpathy). As models do more of the work, reading their output becomes the bottleneck. His ladder from worst to best: plain prose, prose constrained to ASD-STE100 (a controlled English written for aircraft maintenance manuals, one idea per sentence, approved words), diagrams, interactive HTML pages, and custom explainer videos in the style of 3Blue1Brown. Widely reposted in Spanish and Chinese translations within hours.
Claude-shaped science (@AnthropicAI, post). Harvard physicist Matthew Schwartz argues there is an "impedance mismatch" between how scientists want to work and what LLMs do best. Instead of treating Claude as a graduate student, he looked for problems shaped to its strengths and built BootLoops, a toolkit for exact calculations. Claude found the same calculations in ecology and population genetics; the links were correct but unremarkable until domain experts steered them to questions their fields care about.
Industry moves (cluster of 4). US prosecutors say an Earthmade Computer owner smuggled over $300M of export-controlled AI servers to China (@rohanpaul_ai). Google confirmed contact with its four orbital Trillium TPUs launched 10-01 (@rohanpaul_ai). Xiaomi's MiMo-V2.6-Pro ranks #5 among open models in Agent Arena and MiMo-V2.6-Flash sits on its cost frontier at $0.04 per task (@arena). arXiv is limiting how many papers an author can submit as submissions grow exponentially (@sarahookr RT).
Small tools. CopilotKit open-sourced OpenDots, a self-hostable take on OpenAI's Dots that works with any harness (@CopilotKit). Zep's Graphiti frames agent memory as six context types, not just retrieved facts (@akshay_pachaar). Photon raised $4.5M to let agents text users on iMessage and WhatsApp (@zodchiii).
Skip. TESCREAL arguments, political reposts, and the course and webinar promos (Advent of Agents, a 2-hour graph-engineering course repost) carry no new substance.