Summary
The morning slot is thin on research substance and heavy on infrastructure positioning. Curated reposts were empty in the 24-hour window. The strongest signal is the cluster of NVIDIA posts on AI tokenomics (how a model's per-use cost economics determine whether it scales or stalls) coupled with Jensen Huang's claim that AI demand is "going parabolic", paired with @ns123abc reporting that Samsung and SK Hynix are only signing long-term HBM deals with no-waiver clauses, projecting memory shortages until the 2030s. Anthropic news dominated the rest of the AI feed: @bcherny previewed a Claude Code /usage breakdown that itemises tokens by Skill, Agent, MCP, and Plugin (a small but pragmatic feature for plugin-economy debugging) and @ns123abc claimed Anthropic now has more revenue and more accessible compute than OpenAI. Cursor announced doubled usage allowances for new Teams-plan invites for a month. Robert Scoble surfaced Cosmo, a startup building a runtime-generated visual UI layer on top of AI agents (a thin take on "AI is eating the OS"). Nothing in the slot maps directly onto a research paper, so no wiki summary pages are warranted from this slot alone.
Posts
NVIDIA AI tokenomics primer (@nvidia plus thread to 2057505444403241148 and podcast post). NVIDIA's accelerated-computing team broke down four pillars of AI tokenomics: token utility (which model, context length, interactivity needed); token demand (raw volume); token supply (infrastructure given utility and demand); token monetization (cost-based vs value-based pricing). The framing is explicit that "getting the use case right is critical to getting tokenomics right" and that infrastructure choice should be derived backward from the customer use case rather than picked first. This is positioning material for Vera Rubin NVL72 (the next-gen GB300-class rack which NVIDIA is pitching as "agentic AI inference at one-tenth the cost per token") but the structured framework itself is useful for grounding the wiki's ongoing discussion of inference economics versus model capability.
NVIDIA "demand going parabolic" with Dell (@nvidia, article). Jensen Huang joined Michael Dell at Dell Technologies World to unveil updates to the Dell AI Factory with NVIDIA. Numbers in the article: worldwide AI infrastructure spending could reach $3-4 trillion by 2030, token consumption projected to grow 3,400% in the same window, 5,000 enterprises already running AI workloads on Dell AI Factories (Lilly, Samsung, Honeywell named). Vera Rubin NVL72 is pitched at one-tenth the cost per token versus the prior generation, agent sandboxes 50% faster on Vera vs traditional CPUs, enterprise data queries up to 3x faster with Vera CPU. The vocabulary now standard inside NVIDIA pitches ("agent sandbox", "token economics") is a tell that the marketing has caught up with the agentic-AI inference workload, which used to be a research-side framing.
NVIDIA GTC Taipei keynote June 1 (@nvidia, event page). Huang's COMPUTEX 2026 keynote is scheduled for Monday, June 1 at 11 a.m. Taipei time (6 p.m. PT Sunday May 31). Worth watching for any new Rubin variant disclosure, networking SKU (the Multipath Reliable Connection paper cited in the Gmail digest is the kind of thing that gets a Vera-Rubin-class follow-up at Taipei), or hardware roadmap update.
Claude Code /usage breakdown coming (@bcherny). Boris Cherny (Anthropic) previewed the next Claude Code:
/usagewill produce a breakdown of which Skills, Agents, MCPs, and Plugins are consuming tokens. CLI today, Desktop next. Small feature but pragmatic: as the Claude-Code plugin economy gets denser (the @kenhuangus Slash-Command-System chapter in this morning's Gmail starred is at chapter 3 of 10 on the same surface), users need attribution per source. Equivalent in spirit to atopfor token cost.Cursor doubles usage for Teams-plan invites (@mntruell). For the next month, new users invited to a Teams plan get doubled Cursor usage. Plain growth lever. Worth noting alongside the Pragmatic Engineer "Antigravity 2.0 takes 'IDE' out of its new IDE" piece (overwhelmingly negative early feedback on Google's redesigned Antigravity coding IDE due to bugs, poor UX, model support, and Gemini token burn) since Cursor's promotion lands in the same week Google's IDE play stumbles.
Memory long deals signal HBM shortage to 2030s (@ns123abc, SK Hynix bonus tweet). Samsung and SK Hynix are reportedly only accepting long-term memory deals with non-waiver clauses (no escape if pricing collapses), and SK Hynix workers are projected to receive a 900K USD bonus this year. The structural read: HBM supply is sold out years forward at premium prices, and the hyperscalers buying it are locking in capacity rather than spot-buying. This is the supply-side flip of the NVIDIA-pitched "demand going parabolic" framing in the same slot. Memory has not been a tractable bottleneck in this wiki's hardware coverage; if HBM lead times push compute supply to 2030s, that reshapes which inference-efficiency techniques actually matter (anything that reduces HBM bandwidth pressure, KV-cache compression and weight-sharing MoE schemes, gets premium).
Anthropic now has more revenue and more accessible compute than OpenAI (@ns123abc). One-line claim. Aligns with the Decoder reporting (carried in RSS today) that Anthropic is approaching a $559M operating profit on $10.9B Q2 revenue, would be the first profitable AI lab. The Gary Marcus follow-up (today's RSS) notes the profit projection depends on a one-time SpaceX compute discount disclosed in the SpaceX S-1, which may equal or exceed the projected profit, so subsequent quarters' profitability is open. Treat the claim of "more accessible compute than OpenAI" as plausible given the $15B/yr Anthropic-SpaceX compute deal disclosed in the IPO, not yet independently verified.
Cosmo agentic UI layer (@Scobleizer). Scoble surfaced Cosmo, a startup that adds a runtime-generated visual UI layer to AI agents (the user types or speaks and the UI generates in real time on the desktop). Tagline "AI is eating the operating system". This is in the same lineage as Google's Antigravity (struggling, per Pragmatic Engineer) and Anthropic's emerging plugin economy (per @bcherny). The thesis that agent UIs need to be generative rather than fixed is plausible. The interesting research-side question this raises (not addressed in the post) is what the latency budget for runtime UI generation looks like at scale, since a generative UI sits in the same critical path as the model itself.
120 startups launching at Founders Inc San Francisco (@Scobleizer). Soft signal. 120 teams launching at Founders Inc tomorrow evening (May 22 PT). "Robots moving. Agents earning. Hardware working. Software shipping." Scoble notes that 18 is the most he can see in one evening. Worth flagging as ecosystem density signal rather than concrete substance, no specific startups named in the post.
AI-relevant but content-thin tweets (cluster of 3). @Scobleizer "me pretending to do work while my agents run 24/7"; @Tesla Model S X Signature Delivery (not AI substance); @ns123abc "SK Hynix workers get 900K bonus". Mentioned for completeness; no body content to gloss.
Off-topic political and personal feed (cluster of 27). @AustinJustice (4 tweets on Austin crime / ALPR cameras / DA Garza, articles attached but no AI angle); @WHFraudTF (2 tweets, US fraud task force); @brivael (14 tweets in French on economics/politics commentary, no AI research substance); @ns123abc SpaceX Starship V3 unveil plus Nicki Minaj at Starbase plus SpaceX 10 GW Bastrop solar factory (3 tweets, infrastructure not AI). Skip.