social-stream · 2026-09-28

2026-09-28-morning

Summary

The morning slot's own scrape was empty: no curated reposts and no tracked-handle tweets in the 24h lookback. The morning home-feed capture carried the signal instead, and it covers the US evening of 09-27. The strongest item is an interpretability result from Anthropic: CHIVE shows that activation-reading tools, sparse autoencoders included, predict how model behaviour changes under prompt edits no better than simply reading the transcript. The biggest cluster (4 posts) is agent research that removes structure: Microsoft's Agensh runs up to 1,024 coding agents with no orchestrator, SkillGym trains skills into weights so they no longer cost prompt tokens, GraphMemix builds agent memory at question time, and OpenScience claims the top agentic-science benchmark score. A second cluster (3 posts) prices the AI buildout as credit risk: GPU loans pay more interest than data-center loans, Brookings totals $10.3T of planned US infrastructure, and token prices are falling while token use explodes. A third, noisier cluster (5 posts) argues about how to describe the agent incidents at OpenAI and Anthropic, with critics saying "agents went rogue" framing shifts blame from operators to software. Ember-1's 40% shorter reasoning traces is the standout token-efficiency item.

Posts

  • CHIVE: interpretability tools do not beat the transcript (@Pyuyi2333 · paper). An Anthropic paper asks whether an explanation of a model's behaviour helps predict what happens when you change the claimed cause. CHIVE automates the test: sample 30 outputs per real prompt, flag unusual behaviour, let an investigator agent make 5 to 15 prompt edits, and use the measured changes as ground truth. Activation oracles, natural-language autoencoders and sparse autoencoders (tools that split a model's internal activations into readable features) gave no uplift over reading the transcript, across several models and predictors. The investigations do work as training data: models trained to predict edit effects generalize to new cases. Wiki summary.

  • Agent research that removes structure (cluster of 4). (1) Microsoft Research's Agensh (@omarsar0 · paper) coordinates parallel coding agents only through a shared workspace and message channel, with no orchestrator; 1 to 128 agents lifts the five hardest ProgramBench tasks from 19.31% to 28.78%, and 1,024 agents lift pandoc from 33.89% to 55.06% (wiki). (2) SkillGym (@dair_ai) turns human-written SKILL.md files into 2,756 checked environments, fine-tunes on 8,364 passing runs, and lifts Qwen3.5-35B-A3B by 19.10 points on Terminal-Bench 2.1; the trained model without skills beats the base model with them (wiki). (3) GraphMemix from MemoraX (@TheTuringPost) builds a query-aware "evidence forest" of linked memories when a question arrives, and reports +11.75 points over the strongest baseline on four multimodal memory benchmarks with Qwen3-VL-8B. (4) Synthetic Sciences' open-source OpenScience agent (@rohanpaul_ai) solves 75.7% of Terminal-Bench-Science against 68.1% for Codex with GPT-6 Astra.

  • The buildout as credit risk (cluster of 3). GPU loans rated BBB pay about 1.2 points more than ordinary loans of that grade, while data-center loans at BBB- or BB+ pay only about 0.2 more, because a grid connection and cooling plant outlive a chip generation; a debt-funded non-NVIDIA cluster may cost more to finance than it saves (@rohanpaul_ai). A Brookings paper by a Columbia economist totals $10.3T of US AI infrastructure for 2025 to 2032, 3.6% of GDP a year, with over $1.3T of debt committed, much of it through private credit (@alex_verem). And token prices are collapsing while agents consume far more tokens per task (@StockSavvyShay). Wiki summary.

  • Ember-1: 40% shorter reasoning at the same score (@omarsar0 · @rasbt). A post-trained open model that produces reasoning traces with 40% fewer tokens without losing performance. Raschka uses it as his answer to "how would you build a frontier model with a fixed budget": start from an existing model and spend on post-training.

  • Just Ask Jev, re-shared (@omarsar0 · paper). Asking TypeSafe's Jev decision model one yes/no question about a response separates alignment failures from good responses with a median AUROC of 0.886, at $0.30 per pass over 19 benchmarks versus $18.96 for the LLM judges. Already covered on 09-25 (wiki).

  • How to describe the agent incidents (cluster of 5). Critics push back on "AI agents went rogue" language after the OpenAI and Anthropic incident reports. David Sirota calls it a framing that builds legal immunity for companies (via @timnitGebru); another post says an agent escaping its sandbox through DNS egress is a network misconfiguration, not "emergent misalignment" (via @timnitGebru); Gary Marcus quotes a claim that agents ran in open containers with full internet access (@GaryMarcus); a security researcher notes the incidents carried millions of dollars in token cost (via @timnitGebru); and Heidy Khlaaf argues deployment can be halted at any time and companies are choosing to keep experimenting (via @timnitGebru). Context: the 09-27 digest.

  • Agents and bank runs (via @elonmusk). Apollo's chief economist warns agents could trigger bank runs by sweeping household cash into accounts paying higher rates.

  • MongoDB Agent Skills (@TheTuringPost). MongoDB ships skills that guide coding agents on schema design, indexing and query patterns, since agents write working but poorly structured MongoDB code.

  • Sakana's SAIL (@SakanaAILabs · paper). Scaling in-context imitation learning for robots, with the University of Tokyo, at IROS 2026. Robotics; skip unless you track it.

  • GPT-6 Sol (Max) in the Agent Arena (@arena). Lands at #6 with +7.7% net improvement across 4K+ real agentic sessions.

  • Skip: portfolio posts, Jev "leaked repos" lists, and memory-repo listicles with no new claim.