social-stream · 2026-09-28

2026-09-28

Summary

Only the morning slot carried posts, all from the home-feed capture of the US evening of 09-27. The afternoon and evening slots were empty, which is normal for US early morning, and no night slot was written. The strongest post is Anthropic's CHIVE: sparse autoencoders and other activation-reading tools predict the effect of prompt edits no better than reading the transcript, a sharp result for anyone betting on interpretability tooling. The biggest cluster is four agent papers that remove structure, led by Microsoft's Agensh running up to 1,024 coding agents with no orchestrator and SkillGym training skills into weights so they stop costing prompt tokens. The buildout-as-credit-risk cluster is the best hardware signal: GPU loans pay a full point more than data-center loans because chips age faster than the grid connection. The five-post fight over "agents went rogue" framing is mostly commentary, and Ember-1's 40% shorter reasoning traces is the one clean token-efficiency item. The daily digest has the deeper treatment.

Posts

  • CHIVE: interpretability tools do not beat the transcript (@Pyuyi2333 · paper) [morning]. Activation oracles, natural-language autoencoders and sparse autoencoders gave no uplift over reading the transcript when predicting how prompt edits change behaviour. The investigations still work as training data. Wiki summary.
  • Agent research that removes structure (cluster of 4) (@omarsar0 · Agensh · @dair_ai · @TheTuringPost · @rohanpaul_ai) [morning]. Agensh lifts pandoc from 33.89% to 55.06% with 1,024 orchestrator-free agents (wiki), and SkillGym adds 19.10 points on Terminal-Bench 2.1 by training skills into weights (wiki). GraphMemix builds memory at question time (+11.75 points), and OpenScience hits 75.7% on Terminal-Bench-Science.
  • The buildout as credit risk (cluster of 3) (@rohanpaul_ai · @alex_verem · @StockSavvyShay) [morning]. BBB GPU loans pay about 1.2 points over comparable loans versus 0.2 for data-center loans, and Brookings totals $10.3T of US AI infrastructure through 2032. Token prices fall while token use per task explodes. Wiki summary.
  • Ember-1: 40% shorter reasoning at the same score (@omarsar0 · @rasbt) [morning]. A post-trained open model that cuts reasoning tokens 40% without losing performance. Raschka cites it as the fixed-budget path: start from an existing model and spend on post-training.
  • Just Ask Jev, re-shared (@omarsar0 · paper) [morning]. One yes/no question to Jev flags alignment failures at median AUROC 0.886 for $0.30 per pass versus $18.96 for LLM judges. Already covered on 09-25 (wiki).
  • How to describe the agent incidents (cluster of 5) (@timnitGebru · @timnitGebru · @GaryMarcus · @timnitGebru · @timnitGebru) [morning]. Critics argue "agents went rogue" shifts blame from operators, and that DNS-egress sandbox escapes are misconfiguration, not misalignment. Context in the 09-27 digest.
  • Agents and bank runs (via @elonmusk) [morning]. Apollo's chief economist warns agents could sweep household cash into higher-rate accounts and trigger runs.
  • MongoDB Agent Skills (@TheTuringPost) [morning]. Skills that steer coding agents on schema design, indexing and query patterns. Minor.
  • GPT-6 Sol (Max) in the Agent Arena (@arena) [morning]. Lands at #6 with +7.7% net improvement over 4K+ agentic sessions.
  • Sakana's SAIL (@SakanaAILabs · paper) [morning]. Scaling in-context imitation learning for robots. Robotics, skip.
  • Portfolio posts, Jev "leaked repos" lists, memory-repo listicles [morning]. Skip.
  • Afternoon and evening slots [afternoon + evening]. No posts.