Summary
The morning slot is thin on research and heavy on markets and politics, but two curated reposts carry real weight. The strongest signal is a security one: a repost of Carnegie Mellon's "SusVibes" benchmark claims that top coding agents (including SWE-Agent on Claude 4 Sonnet) pass functional tests 61% of the time, yet over 80% of those passing solutions are insecure, a near-total decoupling of "works" from "safe." The second curated repost amplifies SIA, the self-improving-agent loop the wiki already covered on 06-08, where one agent rewrites another's harness or trains its weights from task feedback. On the product side, Cursor shipped Auto-review (a 97%-accurate classifier that grades each agent action allow/block/ask) as the new default, and Goodfire introduced predictive data debugging to surface broken guardrails in training sets before training. The rest is mostly money and noise: a tight cluster around SpaceX's Nasdaq debut and Musk-as-trillionaire, Jeff Bezos's Prometheus closing $12B for an "artificial general engineer," Google's head of Android security resigning over the Pentagon AI deal, and a Grok Build / Terafab product cluster from the xAI orbit. A large volume of US and French political posting (@WHFraudTF, @brivael) is off-topic and skipped.
Posts
SusVibes: vibe-coded agent output passes tests but is insecure (@bayesiansapien repost of @HowToAI_). Carnegie Mellon built a benchmark of 200 real repository-level software-engineering tasks from open-source projects and handed them to the best coding agents on the market. The agents wrote code that passed the functional tests 61% of the time, but over 80% of those functionally-correct solutions contained security vulnerabilities. The framing is pointed: the strongest models are the ones most able to ship working-but-insecure code, because the "tests pass" signal the vibe-coding workflow trusts is blind to security. This is the research-side case for the same day's Cursor Auto-review launch. See wiki summary.
SIA, the self-improving agent loop (@bayesiansapien repost of @rohanpaul_ai). The repost describes SIA: one AI watches how a task agent performs, then either changes the agent's outer setup (prompts, tools, retry rules, output parsing) or trains the model's weights from task feedback. The claim is that the agent improves faster when it can rewrite its own setup and update its model, rather than relying on humans to hand-tune those by hand. This amplifies a paper the wiki already logged on 06-08 as SIA self-improving harness + weights, and it sits squarely in the "manufacture the substrate the agent learns from" thread the digest has tracked all week.
Cursor ships Auto-review as default (@cursor_ai, blog). A classifier subagent now reviews each agent action in context before deciding to allow, block, or ask for approval, reported at 97% accuracy with most misses near ambiguous edges. The blog's framing is that autonomy should be a dial, not a switch: asking for permission too often trains users to rubber-stamp prompts, which is its own safety failure, while too little oversight lets local agents touch credentials and production systems. This is the product mirror of today's SusVibes finding, governing what an agent ships with an external reviewer rather than the agent's own judgment.
Goodfire's predictive data debugging (@eliebakouch reposting @GoodfireAI). Goodfire introduced a way to reveal and shape what a model will learn from a dataset before training. Applied to DPO preference datasets, they report finding broken guardrails, hallucinations, and junk content. HuggingFace's Elie Bakouch endorsed it as "a new way to look at the data." It is an interpretability-for-data-curation play, upstream of the verifier-trust theme running through today's training papers.
Jeff Bezos's Prometheus closes $12B (@ns123abc). The startup raised from JPMorgan, Goldman, BlackRock, and Bezos himself, framed as building an "artificial general engineer." It is the headline funding event of the day and adds another mega-cap entrant to the applied-AI race.
Google's head of Android security resigns over the Pentagon AI deal (@ns123abc, resignation note). René Mayrhofer, Principal Engineer for Android Platform Security, published a farewell note saying "Google management has lost its moral compass," tied to the company's defense work. A responsible-AI / governance signal about internal dissent at a frontier lab.
Grok Build and Terafab cluster (cluster of 4: @JasonBud, @milichab, @ns123abc, @xai). xAI is pushing Grok Build as a general code/design tool: a MongoDB plugin landed in its marketplace for prompt-driven database work and vector search, and several posts show Grok Build generating websites, custom animation tooling, and performant C. Separately the Terafab site (the Tesla/SpaceX/xAI chip-building effort) was rebuilt with Grok Build. Product-velocity signal from the xAI orbit, light on technical substance.
NVIDIA's agentic design workflow (@nvidia). NVIDIA showed an agent (Nous Research's Hermes Agent on RTX Spark) orchestrating Rhino, ComfyUI, and Blender to turn a concept sketch into a photoreal render in one workflow. The pitch is that engineering and design workflows are increasingly defined by the connections between tools, not the tools themselves, an agentic-orchestration framing aimed at creative professionals.
Kilo Code Console beta and Fable 5 plan change (cluster of 2: @kilocode, @kilocode). Kilo launched a local browser-based UI for managing projects, git worktrees, sessions, and settings (a face for the CLI, no more hand-editing JSON). The more newsworthy post: Anthropic's Fable 5 comes off Pro, Max, Team, and seat-based Enterprise plans on June 23 and moves to usage credits until capacity allows, while Kilo's gateway keeps it at $10/$50 per million tokens. A concrete data point in the inference price war.
SpaceX IPO and trillionaire cluster (cluster of 3: @brivael, @ns123abc, @PeterDiamandis). SpaceX began trading on Nasdaq, prompting a wave of posts that Founders Fund turned $600M into $50B and that the IPO makes Musk the world's first trillionaire. Markets-and-hype, not AI substance, but high-volume on the feed today.
US and French political posting (@WHFraudTF, @brivael). A large block of LAHSA-fraud and French-politics content. Skip, no AI relevance.