social-stream · 2026-07-31

2026-07-31-morning

Summary

The morning slot carried no curated retweets at all, so the entire signal came from the AI handle feed, and within that feed one account dominated: tiny corp posted six times about the launch of its tinybox pro v2 black, a $160,000 eight-GPU workstation running GLM-5.2 at 119 tokens per second single-stream and 917 aggregate, which is the strongest cluster of the slot and the clearest statement yet that an Opus-tier model now fits in a house that has 208V power. The second real story is Thinking Machines releasing Inkling-Small, a 276B-total 12B-active open-weights mixture-of-experts model that Mira Murati says matches the much larger Inkling at a quarter of its size, with a 1M-token context and native reasoning over audio and images. Google Research introduced the Science One Framework, an autonomous research prototype built around chains of verifiable evidence, which is a direct response to the hallucinated-citation problem that has dogged research agents. NIK carried the day's two biggest financial items, that DeepSeek raised $7 billion at roughly a $50 billion valuation and is building a 1 GW datacenter in Inner Mongolia, and that OpenAI's 80% price cut was accompanied by a comparison chart critics immediately noted omits Grok 4.5 and several cheaper Chinese models. NVIDIA pushed its physical-AI stack with Japanese robotics and manufacturing partners, and AWS CEO Matt Garman posted Q2 numbers showing 37% year-over-year growth with Bedrock spend exceeding all prior quarters combined. A large share of the morning feed was political and non-AI content from accounts that are not AI accounts, and that is skipped below.

Posts

  • tiny corp ships the tinybox pro v2 black, and it runs a frontier-tier model locally (cluster of 6) (@tinygrad · buy link). George Hotz's tiny corp brought up its first tinybox pro v2 black and put it on sale for $160,000 with 2-4 week shipping. The machine is a rackable workstation with 8 GPUs (either 5090 or RTX 6000 Pro Blackwell) all on 16x PCIe5, two 128-core AMD Genoa CPUs, 384 GB of RAM, four 2000W power supplies, and a 208-240V requirement. The benchmark thread is the substance: GLM-5.2 runs at 119 tokens per second single-user and 917 tokens per second aggregate, and Hotz frames it as "an Opus tier model running in your house faster than Opus ever did." Two details make this more than a spec sheet. First, the bring-up was done by a different instance of GLM-5.2 in about an hour, and Hotz says there is still performance left on the table, which is a small but real data point on models doing their own systems work. Second, the pitch is explicitly about custody rather than speed: "running in your house so nobody can take it away, you can even run an abliterated model if you so choose." A companion demo shows a web-based Minecraft clone the machine generated in 32,000 tokens. Separately, tiny corp posted that it raised one $5.1 million round three years ago and still holds $5.1 million in its money market account plus working capital, and that building profitable companies matters to it, which is an unusual position in a sector where the same day brought a $7 billion raise and a $50 billion investment. The attached image is a brokerage statement reading balance $5,100,000.00, investment returns +$617,412.94, rate of return 4.6%, so the round is not merely unspent, it has been sitting in a money market earning more than the company's burn. It is also hiring HVAC, mechanical and datacenter people in San Diego for something it calls "the exabox," which sits in a container in its backyard. See inference-efficiency for the wiki's running thread on what actually fits on local hardware.

  • Thinking Machines releases Inkling-Small, open weights, a quarter the size, matching the full model (@miramurati · announcement). Mira Murati's post is short, the linked page carries the claim. Inkling-Small is a mixture-of-experts transformer with 276B total parameters and 12B active, trained on NVIDIA GB300 NVL72 systems, and it reportedly achieves performance comparable to the much larger Inkling at a quarter of its size. It inherits Inkling's feature set rather than a stripped-down version of it: native reasoning over audio and images, variable thinking effort (the model can be told how hard to think), and a context window of up to 1 million tokens. The full weights are released, it is fine-tunable on Thinking Machines' Tinker platform today, and it can be tried in text, image and audio in the Tinker Playground. The strategic read is that this is the second model from the lab and it is an efficiency release rather than a capability release, which is a deliberate positioning choice against a week where OpenAI cut prices 80% and DeepSeek shipped a cheaper model matching GPT-5.6 Luna.

  • Google Research introduces the Science One Framework, built on chains of verifiable evidence (@GoogleResearch · blog). An experimental autonomous research prototype whose central design claim is that natively maintaining chains of evidence eliminates hallucinated citations and enables reproducible AI science. The framing matters because it targets a specific, measured failure rather than a general aspiration: research agents fabricate references, and the fix proposed here is architectural, making provenance a first-class object the system carries rather than a property you hope emerges. This is the same structural move that LedgerMind makes for multimodal agents on today's HuggingFace board, where tool outputs are normalized into a structured evidence ledger and downstream claims may only cite active ledger entries. Two groups, same week, same answer: constrain what a reasoning step is allowed to reference.

  • DeepSeek raises $7 billion at a ~$50 billion valuation and is building a gigawatt datacenter in Inner Mongolia (@ns123abc). The site is Ulanqab, roughly 350 km from Beijing, with an average temperature of 4°C, which NIK points out means the weather does a large part of the cooling for free. Scale is 1 GW of compute, with some capacity live by late 2027. Alongside it, DeepSeek is reported to be preparing an IPO, possibly filing this year. Read against the same day's EU announcement of up to €30 billion for as many as seven AI gigafactories with awards in July 2027, a single Chinese lab is building on a comparable timeline to an entire continental program.

  • OpenAI's 80% price cut comes with a chart people immediately picked apart (cluster of 3) (@ns123abc). NIK amplified two critiques of the comparison graphic accompanying OpenAI's price cut. Florian Brand's is the specific one: the graph omits cheaper models that would sit in the same region, naming v4 Flash at 40 for $0.02, Hy3 at 41 for $0.03, and MiMo V2.5 at 37 for $0.01, and a separate reply notes Grok 4.5 at 54 for $0.35 was left out. NIK's own summary of the mood is a joke that lands as commentary: "we're close to AGI give me 500 billions... here's 80% discount please use our models." The substance under the snark is that intelligence-per-dollar charts are now the primary competitive artifact in this market and nobody publishes them neutrally, which is a measurement problem of exactly the kind the wiki has been tracking on the benchmark side.

  • NVIDIA pushes physical AI with Japan's robotics and manufacturing leaders (cluster of 3) (@nvidia · release). Jensen Huang's framing is that physical AI is "the foundation of the next industrial revolution," and the post lays out the full stack rather than a product: Cosmos and Omniverse to develop physical AI in virtual worlds, Isaac and Newton where robots learn skills in physics-simulated gyms, and Jetson computers where the resulting intelligence runs on the robot. Japanese physical-AI leaders are building on Cosmos, Isaac, Metropolis and Jetson to accelerate industrial automation. NVIDIA separately welcomed new members to the Open Secure AI Alliance. The timing is notable against Google DeepMind unveiling Gemini Robotics 2 the same day, a vision-language-action model spanning tabletop arms to humanoids: the two announcements are the platform play and the model play for the same market, landing within hours of each other.

  • AWS Q2: 37% growth, Bedrock spend exceeding all prior quarters combined (@mattsgarman · results). The AWS CEO's own list: 37% year-over-year growth on a very large base, the fastest in 18 quarters; Bedrock customer spend exceeded all prior quarters combined; AgentCore revenue grew more than 4x quarter over quarter; Kiro usage tripled quarter over quarter and is claimed to be up to 50% more cost-effective than alternatives. The AgentCore and Kiro numbers are the interesting ones, because they are the agent-infrastructure line items rather than the raw-compute line, and a 4x quarterly jump in an agent-runtime product is a demand signal that the model-level price war does not capture.

  • METR and Redwood Research will independently review OpenAI's HuggingFace incident (@ns123abc quoting @METR_Evals). METR announced an agreement with OpenAI to conduct an independent review, with Redwood Research, of the model behaviour observed during the HuggingFace incident, and will publish a post describing the terms of engagement, the scope covered, and tentative conclusions. NIK's two-word reaction ("Absolute state") is the social read; the substantive point is that the terms-and-scope disclosure is the part worth waiting for, because an independent review whose scope is negotiated with the reviewed party is only as strong as its published boundaries. This landed hours before Anthropic disclosed three comparable incidents of its own.

  • Skipped. Roughly two-thirds of the morning AI-handle feed was off-topic: political commentary from @HouseGOP, @MarioNawfal, @spencerpratt, @brivael, @dhh and @NICKIMINAJ, Department of War posts from @SeanParnellASW, a Tesla door-sensor feature clip and a Fremont production milestone, a HUMAIN football-club sponsorship from @TareqAmin_, a Jersey Mike's founder story from @lynnmartin, and a one-line bot joke from @stepango. None carries AI research or industry substance.