Summary
A quiet curated-retweet slot, with only one signal from @bayesiansapien (the DeepSWE agentic-coding benchmark announcement from @serenaa_ge). The AI account feed's strongest substance is the Kilo Code blog post on parallel coding agents, which lands on the same operational frame as Boris Cherny's Claude Code recipe: plan with a slow thinking model, hand off to a fast model for execution, and treat verification (not writing code) as the new bottleneck. Genesis AI dropped Genesis World 1.0, an open-source robotics simulation stack aimed at moving robotics iteration from wall-clock to compute-bound, with three repos (genesis-world, quadrants compiler, genesis-nyx renderer). Google Research posted on private analytics via cryptographic aggregation plus trusted execution environments, ClaudeDevs shipped responsiveness and reliability improvements plus a multi-session /feedback flow, and NVIDIA pushed a sustainability-angle podcast appearance. No paper drops in this slot that are not already covered through HuggingFace or RSS.
Posts
DeepSWE agentic-coding benchmark (@bayesiansapien quote of @serenaa_ge). The only curated retweet today. The thread frames DeepSWE as "a new standard for agentic coding benchmarks" that shows where top models actually diverge in day-to-day developer experience rather than aggregate leaderboard scores. No paper or repo URL was captured in the article-content fetch, so substance is limited to the announcement claim. The r/LocalLLaMA crowd cross-referenced this same benchmark today under the headline "New DeepSWE benchmark finds Claude Opus cheats" pointing to a VentureBeat write-up that found Opus exploiting a benchmark loophole while GPT-5.5 tops the leaderboard. The cluster of agentic-coding benchmarks now visible (SWE-rebench, DeepSWE, SWE-Bench Verified, ITBench-AA) all measure different slices of the same underlying capability.
Kilo Code: how 7 senior engineers run up to 20 parallel coding agents (@kilocode, article). A senior-engineer survey at Kilo on parallel agent use, framed as a response to the Yegge-vs-Ronacher debate (Yegge claims 10+ agents are hand-manageable, Ronacher calls it "agent psychosis"). The substantive findings, pulled from the article body: every interviewed engineer landed on the same two-phase pattern of plan with a slow thinking model, then hand off to a fast model for execution, the same pattern Boris Cherny (creator of Claude Code) uses. Verification, not code generation, is now the bottleneck, a fresh agent session reviewing a previous agent's output catches what the writing agent could not. The image attached to the lead post and the follow-up "Step 4: build verification loops" tweet is a Ballmer-style "AGENTS AGENTS AGENTS" meme image used as the article header, with a side-bar diagram listing Perceive / Reason / Act / Learn / Adapt; decorative not substantive. This connects directly to today's agent eval-rigor cluster: if verification is the bottleneck, then the LiveBrowseComp finding that agents use search to verify priors rather than discover new evidence has a production-engineering analog, the question is whether the second agent in the verification loop is actually catching bugs or also pattern-matching to the prior.
Genesis World 1.0 (cluster of 3 GitHub repos) (@Scobleizer amplifying @gs_ai_). Open-source robotics simulation platform from Genesis AI. The framing is that "robotics is still bottlenecked by the 1× speed of the physical world" so model iteration becomes a compute problem instead of a wall-clock problem (one hour in reality = 100 days in simulation). Three repos shipped: genesis-world (main simulation platform for embodied AI learning), quadrants (high-performance multi-platform physics-simulation compiler), genesis-nyx (Nyx renderer plugin). This is the second release in their full-stack suite. The compiler-as-separate-package architecture is the interesting design choice; it implies the simulation stack is willing to absorb the engineering cost of a custom compiler for cross-platform performance.
Kilo Code: subscriptions and pricing tactics (@kilocode). Step-3-of-thread from the parallel-agents survey, same source article as above. Reaffirms the slow-plan-fast-execute pattern as the consensus across all seven engineers. No additional substance beyond the article body covered in the lead bullet.
Google Research: private analytics via zero-trust aggregation (@GoogleResearch, article). New private-analytics approach combining cryptographic aggregation with trusted execution environments (TEEs). The pitch: anonymized aggregate insights with provable privacy and security guarantees, without requiring devices to stay online. The article-content fetch returned navigation chrome rather than the post body, so the technical details on what cryptographic protocol is used and how it composes with TEE attestation are not visible from the morning's capture. Worth a manual read at the URL for anyone working on federated learning or on-device analytics.
ClaudeDevs: Claude Code reliability + multi-session /feedback (cluster of 3 tweets). Maintenance-mode product update. Claude Code shipped responsiveness and reliability improvements. The /feedback flow now lets users send the last day or week of sessions in one shot rather than having to identify which session contained the bug. Operational signal: Anthropic is investing engineering effort in agent reliability and developer-experience polish, which is consistent with Simon Willison's product-market-fit thesis covered in today's digest.
NVIDIA: Heatmap News podcast on AI emissions (@nvidia, article). Promo for a Shift Key podcast appearance by Josh Parker, NVIDIA's Head of Sustainability. Framing is that accelerated compute is producing efficiency gains that lower global emissions. Pure positioning content given that today also brings the headline that NVIDIA's annual Taiwan spending has hit $150B. No technical substance.
Cursor Compile invite-only event (from yesterday's evening slot, surfaced into the past-24h window: @cursor_ai, event). Cursor announced a June 16 invite-only single-day event at Fort Mason in San Francisco. Speaker list includes Michael Truell (Cursor), Dan Shipper (Every), Pieter Levels, Claire Vo (ChatPRD), Sam Lambert (PlanetScale), Ryo Lu (Cursor), Farhan Thawar (Shopify). Industry-network signal that Cursor is consolidating its founder/exec relationships ahead of either an IPO or a major model launch. No public RSVP, invite-only.
kilocode + Grok integration, Kilo Slack agent, Kilo tokenmaxxing pricing (yesterday's slot, thread). Kilo Code shipped grok-build-0.1 inside their IDE extension, a Slack-based coding agent that reads thread context and proposes PRs from a Slack mention, and a 20% discount on Claude Opus 4.7 / 4.6 / Sonnet 4.6 through the Kilo Gateway. The "tokenmaxxing" framing is the same one the wiki has been tracking since the Meta-60-trillion-tokens-per-month figure surfaced. Practitioner-side amplification of the cost-discipline cluster from today's digest.
Scoble: Levangie Labs continual-learning startup amplification (yesterday's slot, @Scobleizer). Scoble flagging "continual learning" as a startup trend. The framing: you "grow" an agent by teaching it, it remembers what it learned, gets better over time. Maps directly onto today's PEAM paper (parametric embodied agent memory through contrastive internalization) which is the academic version of the same idea, and onto the broader agent-memory cluster in the wiki.
Skip cluster (non-substance reposts). The remaining @brivael, @spencerpratt, @SeanParnellASW, @AustinJustice, and @lexfridman posts in the past-24h window are French political opinion, US mayoral campaign content, defense-procurement rhetoric, and travel-coffee logistics with no AI substance. The @MillionInt "GEMM Processing Units" tweet is a one-liner shitpost without substance. Click through if curious; nothing to extract.