Summary
Today was NVIDIA's day across both slots. The strongest cross-slot cluster is the GTC Taipei keynote (Jensen Huang in Taiwan), which surfaced morning and evening with roughly fifteen posts covering the RTX Spark consumer agent PC, the Vera "CPU for Agents," the 550B Nemotron 3 Ultra, Cosmos 3 for physical AI, and the MGX AI-factory platform on Vera Rubin. Captured keynote slides give hard specs: RTX Spark is a 1-petaflop FP4 Blackwell GPU with a 20-core MediaTek-built Grace CPU and 128GB unified memory, and Vera is an 88-core agent-tuned CPU claiming 40% lower loaded latency than x86. The second real thread is agent security, all in curated reposts on the same day: Tsinghua's Chain-of-Authorization (claiming a 98.5% attack success rate cut to 0%), Anthropic's Zero Trust for AI Agents framework, and NVIDIA's SkillSpector skill scanner. The single strongest non-NVIDIA standout is the evening's MiniMax M3 open-weights launch, with day-zero benchmarks (59.0% SWE-Bench Pro, 66.0% Terminal Bench 2.1) and sparse attention scaling to 1M tokens, extending the M2-Mini work already in the wiki. Smaller but worthwhile signals include the harness-beats-model "binding constraint" thesis, a MIT/Stanford/NYU/Princeton paper arguing AI's productivity feeling can outrun reality, and the Guild local-memory tool. A large block of French political and Los Angeles crime posts in the account feed carried no AI substance and is skipped.
Posts
- NVIDIA GTC Taipei keynote: RTX Spark, Vera, Nemotron 3 Ultra, Cosmos 3, MGX (cluster of ~15) (@nvidia, @Scobleizer, @ns123abc · NVIDIA Newsroom) [morning + evening]. Jensen Huang's keynote dominated the feed all day. RTX Spark is pitched as the first Windows PC built for personal AI agents: a 1-petaflop FP4 Blackwell GPU, a 20-core MediaTek-built Grace CPU, and 128GB unified memory at 600 GB/s. Vera, titled "CPU for Agents," claims 88 cores, 40% lower loaded latency than x86, and 3.4 TB/s core-to-core bandwidth tuned for latency-bound orchestration. Nemotron 3 Ultra (550B) tops long context (95% on Ruler at 1M tokens) but trails GLM on planning and Kimi on coding. The evening added the MGX AI-factory platform on Vera Rubin with 800 VDC power. Full writeup: NVIDIA GTC Taipei summary.
- MiniMax M3, first open-weights model to combine three frontier capabilities (@kilocode · MiniMax docs) [evening]. Day-0 launch: 59.0% SWE-Bench Pro, 66.0% Terminal Bench 2.1, 74.2% MCP Atlas, natively multimodal, with MiniMax Sparse Attention pushing context to 1M tokens. Sparse-attention-at-scale plus open weights makes this the day's only non-NVIDIA must-read; extends MiniMax M2-Mini.
- Chain-of-Authorization bakes permissions into the reasoning chain (@alex_prompter) [morning]. A repost of Tsinghua work arguing every enterprise agent shares one flaw: the LLM treats all accessible data as fair to share, and the permission layer is "just a prompt." Baking authorization into the chain-of-thought reportedly cuts attack success from 98.5% to 0%. Tweet-level evidence, no paper text. Wiki: Chain-of-Authorization.
- Anthropic's Zero Trust for AI Agents framework (@_vmlops · PDF) [morning]. A zero-trust playbook naming threat classes traditional access controls miss: prompt injection through external data, tool poisoning via MCP metadata, memory-based privilege retention, and multi-agent pivot attacks. Defenses split into Foundation, Enterprise, and Advanced tiers. Pairs with Chain-of-Authorization and ClawTrojan.
- NVIDIA SkillSpector: a security scanner for agent skills (@bibryam · GitHub) [morning]. Open-source tool that scans agent skills before install, running 64 checks across 16 categories (prompt-injection, credential-theft, supply-chain, AST plus taint-flow) with an optional LLM review layer and SARIF output. The tooling complement to the day's agent-security cluster, shipped the same week NVIDIA ships agents to consumer machines.
- Nous Research Hermes runs on RTX Spark (@ns123abc) [morning]. Hermes Agent now runs on RTX Spark via the new OpenShell runtime, connecting it to Microsoft's security primitives. The substance behind the broader Hermes promotion: the open agent framework is now part of NVIDIA's consumer agent stack.
- The binding constraint thesis: harness beats model (@IntuitMachine) [morning]. Argues "better model equals better agent" is wrong: GPT-5 still fails 60% of long coding tasks, but the same weights with a better harness deliver 10x. An agent's ceiling equals min(model, harness), and the harness is currently the binding constraint. The conceptual opposite of the "compile the workflow into weights" paper in today's digest.
- MIT/Stanford/NYU/Princeton: AI's productivity feeling can outrun the reality (@rohanpaul_ai) [morning]. A paper finding AI makes people feel more efficient even when the measured benefit is tiny or negative. The danger is a feedback loop: once people use AI they reach for it again, even on easy tasks where doing it themselves would be just as fast.
- Guild: shared local memory for multi-agent coding (cluster of 2) (@AlphaSignalAI · GitHub) [morning]. A single Go binary backed by embedded SQLite running a local MCP server so Claude Code, Cursor, and Codex share memory and hand off work, nothing leaving the machine. Search fuses BM25 keyword matching with vector similarity.
- Hermes Agent Control Room and masterclass promotion (cluster of 3) (@AlphaSignalAI, @cyrilXBT) [morning]. Reposts promoting the Hermes framework: the Control Room template treats an agent fleet like an OS growing in four stages, plus a "zero to production agent in one sitting" masterclass. X article bodies unfetchable; promotional rather than a research result.
- NYSE exploring NVIDIA Vera CPUs (@lynnmartin) [evening]. NYSE got a keynote shout-out and is looking at Vera CPUs to scale capacity and cut latency for market infrastructure. A concrete enterprise pull-through for the Vera Rubin platform.
- Claude removed temperature controls, and so did everyone else (@alex_prompter) [morning]. Pushback on the reaction to Claude removing manual temperature controls, noting OpenAI locked it on o1/o3/GPT-5 and Google warned against changing it on Gemini 3. The argument: forcing temperature to zero collapses reasoning models' multi-pass verification into one greedy line. Opinion, not a result.
- Kilo Code: Opus 4.8 at 4.7 pricing, Step 3.7 Flash free (cluster of 2) (@kilocode · blog) [morning + evening]. Opus 4.8 at the same price as 4.7 (evening: 20% off in Kilo, quoting Anthropic's honesty framing), StepFun's Step 3.7 Flash free, and Xiaomi cutting MiMo pricing by up to 99%. A concrete data point on model pricing pressure.
- AI Layoff Trap economics paper (@jackcoder0) [morning]. A repost dramatizing a March 2026 Wharton and Boston University paper modeling automation-driven demand collapse: an economy that produces everything and sells to nobody. Rhetorical post over a peer-reviewed macro model, the economic counterweight to the day's hardware optimism.
- GPT-Realtime 2.0 voice computer control demo (@Scobleizer) [morning]. A reshared demo of controlling a computer entirely by voice via GPT-Realtime 2.0. A demo video, no benchmark, but notable alongside the NVIDIA agent-PC push.
- Opaque x.com article reposts (@cyrilXBT, @zodchiii, @ashwingop, @dair_ai) [morning]. Four curated reposts to X native long-form articles whose bodies could not be fetched (expired cookies). The dair_ai one is the DAIR.AI Top Papers of the Week roundup, recovered separately from Gmail and fed today's digest.
- Political and crime account-feed posts [morning + evening]. Skip. A large block from @brivael (French election commentary, ~18 posts), @spencerpratt (LA politics), @AustinJustice, @WHFraudTF, and @lynnmartin's feed neighbors carried no AI substance.