social-stream · 2026-08-31

2026-08-31-morning

Summary

The morning slot is empty, and for the fifth day in a row the reason is infrastructure rather than a quiet field. The Twitter farmer ran at 09:09 IST with a 24-hour lookback and returned zero curated @bayesiansapien reposts, zero AI-handle tweets and zero enriched articles, because no Nitter instance was reachable; the preceding evening run at 20:50 IST on 08-30 returned exactly the same empty result set, so there is no overnight tail to recover either. The saved-posts channel is the healthier one and it tells a different story. Bookmarks come through X's GraphQL API on session cookies rather than through Nitter, that path authenticated and read the timeline normally, and it returned zero newly-saved posts for the second consecutive day. That is a measured zero, not a failed read, and the two silences should not be summed into one signal: the public scrape is a broken pipe and the empty bookmark delta is a real observation that nothing was saved. Where a working feed would have carried today's cluster, the same material arrived through starred Gmail instead, and it was unusually dense: DAIR.AI's weekly roundup landed six substantive items, five of which are the same harness-and-context story the saved reading trail has been walking since 08-25. The strongest single number in the whole morning came from that roundup rather than from any tweet, and it is Prime Intellect's open-source harness moving ARC-AGI-3 Best@1 from 30% to 95.5% with the model class held fixed.

Posts

  • No posts captured. Both the curated repost feed and the 71-handle AI account feed returned nothing, in both the 08-31 morning run and the 08-30 evening run before it, because no Nitter instance was reachable. Nothing was dropped by a keyword filter and nothing was skipped for relevance. There is no partial capture to report and no handle-level detail to give, because the fetch never reached a server. This is the fifth consecutive slot with the same cause.

  • No new saved posts. The bookmarks channel authenticated successfully against X's GraphQL Bookmarks operation and read the full timeline, so this is a genuine zero. The running theme tally is unchanged: loop and harness engineering at 13 saves and still dominant, inference and KV cache and GPU kernels at 5. The trail's most recent entry is still the 08-29 save separating the four cache layers, whose sharpest finding was that provider prompt-cache entries are keyed to a model, so routing mid-session to a cheaper model reprices the whole accumulated conversation at cold rates.

  • The harness cluster that would have owned a working feed arrived through Gmail instead (cluster of 5) (DAIR.AI weekly roundup). Five of its six items are one story about where agent capability actually lives, and read together they are the sharpest statement of the saved trail's dominant theme this month. Prime Agent, open-sourced by Prime Intellect, persists histories, memories, skills, prompts and subagent specifications across trajectories rather than resetting everything but the files on disk, and gives the model a persistent IPython session so it processes its context programmatically instead of re-reading a flat transcript; holding the model class fixed, ARC-AGI-3 Best@1 goes from 30% to 95.5%. Scroll, from Alibaba, removes the memory schema entirely: each session is backed by an append-only event log plus a sandboxed Python kernel, tool outputs bind to typed variables across model calls rather than being re-serialized into the prompt every turn, and only explicitly printed projections cross into the working view, reaching 94.8% on LongMemEval_S and 73.1% on BEAM_10M, 5.1 points over the best published memory system. JIT-Agent goes one level up and makes the harness itself the model's output, synthesized per task under a four-module protocol covering memory, planning, action protocol and tool orchestration, patched mid-run; DeepSeek-V4-Flash passes GPT-5.6 on DeepSearchQA by 9.1 points and GLM-5.2 gains up to 20.2. NVIDIA's Skill Lift is the measurement counterweight: across 145 real skills from internal and public catalogs, the structural scanner most enterprises gate on correlates with actual judged quality at a Spearman rho of 0.14, meaning the review gate tells you a skill is well formatted and nothing else, so ACES proposes running the same task twice under identical model and sandbox and scorer, once with the skill loaded and once without, and measuring the difference. And Netflix's judge-lifecycle writeup is the rare LLM-judge paper with production consequences attached, running judges over hundreds of thousands of recommendation explanations weekly across four phases (birth, training, deployment, monitoring) with a meta-judge over the judge's own reasoning as the learning signal, validated by a five-week A/B test over tens of millions of members. The through-line is the one agent harness engineering already holds: capability lives in the scaffold, not only in the weights.

  • The paper feed came back and it converged on one idea (no social signal, noted here because the morning's substance is entirely elsewhere). HuggingFace rolled its daily-papers date forward for the first time since 08-28 with twelve papers, and three of them plus a Kurate leaderboard entry make the same argument about training cost: a single scalar reward applied uniformly across a whole rollout discards structure the reward already contained. ContextPilot localizes the advantage to individual context-editing actions, RCCA to the code span a rubric complaint was about, and StepGuard's Balance-GRPO to the safe-versus-unsafe action class. Full treatment in the 08-31 digest.

  • Skip. Nothing to filter. There were no promotional posts, no job ads and no opaque x.com/i/article/ reposts to group, because there was no feed at all.