social-stream · 2026-10-01

2026-10-01-afternoon

Summary

This slot had no curated reposts. All of its signal comes from the 139-post Following-feed capture, and roughly half of that is promoted ads and off-topic noise. The strongest item for this reader is DeepSeek open-sourcing its TileLang kernel stack for Huawei Ascend, with matmul kernels at 99.8% of peak. That is a real step toward Nvidia-free training, and it arrives in the same quarter Huawei says it will ship training chips. The second signal is a compute-deal cluster (cluster of 5): Oracle selling Tencent $7B of chips, Ed Zitron's back-of-envelope that the deal prices out at about $1.60 per GPU-hour, Dell's $15B Tokyo site, and SpaceX scaling to 2.3 GW. On the agent side, Meta's Meta-Reasoning Agent result fits the harness thread cleanly: a separate controller that decides how to spend compute keeps scaling after a plain agent stalls. Also worth a click: Perplexity's contextual embedding model and a clear explainer of the LLM serving scheduler. Skip the Grok and Grokipedia posts, the political threads and the ads.

Posts

  • DeepSeek ports its kernel stack to Huawei Ascend (@rohanpaul_ai · TileKernels repo). DeepSeek released six Ascend projects. The core one is TileLang, the kernel language behind most of V4's training operators, so the same source now compiles for Huawei chips. DeepGEMM-Ascend reports 431 of a possible 432 BF16 TFLOPS. DeepEP-Ascend shipped without a license file, so nobody can legally reuse it yet. Related: DeepSeek V4 on Ascend.

  • Compute deals and their unit economics (cluster of 5). (1) Oracle signed a $7B deal to supply Tencent with about 100K advanced chips in Southeast Asia, and Tencent is paying about 30% upfront (@StockSavvyShay). (2) Ed Zitron worked the deal through: assuming a five-year term, it comes to about $1.60 per GPU-hour, cheap for an H100 and very cheap for Blackwell (@edzitron). The term and chip type are guesses, but this is the right question to ask. (3) Dell is leading a $15B, 400 MW data center near Tokyo, part of Japan's plan for up to 4 GW over five years (@StockSavvyShay). (4) SpaceX is heading for about 2.3 GW by end of November, adding 220K GB300s this week and another 220K in November, roughly 440 MW per batch (@rohanpaul_ai). (5) Musk reposted "Starmind," a pitch for orbital data centers, as the way around permits and power. Treat it as a pitch (@elonmusk RT). On the financing side, see AI buildout financing risk.

  • Meta: a manager agent makes extra compute pay off (@rohanpaul_ai · arXiv 2609.38147). "Thinking Before Thinking" splits the agent in two. Workers do the task. A controller keeps a progress summary, weighs options against the remaining budget, and decides which past results each worker sees. At the largest budget it won all 12 head-to-head tests. With GPT-5.5 on a coding benchmark, tripling the budget took it from 64.1% to 71.5%, while the baseline stalled near 64%. This is budget allocation as a harness component. See agent harness engineering.

  • The LLM serving scheduler, explained (@_avichawla). A clear walkthrough of one request's path through a serving engine. Every step, the scheduler checks the token budget and free KV-cache blocks, puts running decodes ahead of new admissions, and chunks long prefills. It mixes prefill and decode in one batch and releases a finished request's cache blocks right away so another request can join (continuous batching). Good primer material, nothing new.

  • Perplexity's contextual embeddings: pplx-embed-v2-context-9b (@denisyarats). The model encodes the whole document once and then pools vectors per chunk, so each chunk sees the full document. Instead of training on a single gold chunk, it distills relevance from Perplexity's context-compression model, which scores every token against the query. It reports +14.4 answer recall@10 over voyage-context-4 on a blind private benchmark. At int8 with 1024 dims (1 KB per vector), it still beats voyage at 8 KB, which is an 8x cut in storage.

  • Looped Diffusion Transformer (@arXivBangers · paper). Looped-DiT runs shared Transformer blocks several times within each denoising step. A 260M model beats a 6.5x larger model on text-to-image benchmarks with 4.9x less inference compute. Weight sharing through depth recurrence is the efficiency idea here. The image domain is secondary.

  • RRSI resurfaces (@rohanpaul_ai · repo). Google's framework evolves the prompts, control flow, tools and memory of an agent harness around a frozen LLM. This is a re-share of a September release; covered at RRSI.

  • Gemini 4 Argon afterglow (cluster of 3). One post claimed Argon is an "RSI-developed" model, and the same author then said the source post was deleted. Treat that as rumor (@IntuitMachine, follow-up). Dan Roy said he contributed to Argon's improved math (@roydanroy). See Gemini 4 Argon.

  • Agents can refine benchmarks too (@tli104). Work led by Seungone Kim iterates on a benchmark with agents instead of iterating on methods. It finds that human oversight still decides benchmark quality. No paper link is in the post, so click through to read.

  • What building RL environments actually involves (@Sanyam0605). Notes from interviews at RL-environment startups name three hard problems: writing fair tasks, making sure a task measures the capability it claims to, and turning tasks into synthetic training data. Useful framing, light on substance.

  • Altman on "ultra fast" prompting (@rohanpaul_ai). Altman says he now prompts only on OpenAI's low-latency tier because a fast feedback loop changes how he thinks. It doubles as a pitch for the Ultrafast tier from DevDay.

  • JetBrains Junie Plan Mode (@jetbrains). Promoted post. A frontier model writes the plan and a cheaper model executes it, for about 40% lower cost. It is a shipped planner/executor routing split, which makes it worth one line.

  • OpenAI distillation accusation (repost) (@rohanpaul_ai). Re-share of the morning story, already covered in the morning slot.

  • AI and jobs policy (cluster of 2). A DeepMind, Oxford and Chicago paper argues that no single policy holds up across scenarios. It proposes expanding unemployment benefits and the low-wage tax credit now, turning the credit into an income floor if joblessness persists, and creating a public investment fund if labor's share keeps falling (@rohanpaul_ai · SSRN). Bill Gates says a robot that replaces a worker should pay that worker's FICA tax (@rohanpaul_ai).

  • Smaller items. TRL telemetry shows 1M fine-tuning runs a month (@huggingface RT). SemiAnalysis noted ICLR submissions have quadrupled since 2023, from 4,938 to 19,525 (@dylan522p RT). MIT used AI to design an mRNA vaccine formulation that keeps at room temperature for a year (@rohanpaul_ai). LiteLLM launched Lens for enterprise AI traffic (@ycombinator RT). Databricks launched AI Decide, a low-cost text-to-decision API (@jjanezhang).

  • Grok Bot and Grokipedia v0.3 (@elonmusk). Launch hype, plus a repost claiming Grok Bot builds itself. No details. Skip.

  • Safety-community politics (@timnitGebru, @pmddomingos). Gebru's threads attack Sanders' AI-safety advisers, and Domingos posts quips. Commentary, no claims to evaluate. Skip.

  • Ads and off-topic (Explee, VPS, proxies, trading apps, robot-gadget hashtag posts, P99 CONF). Promo and filler. Skip.