hardware · 2026-09-23 · Tier 1

Semiconductor Week 38: memory crosses half the market, and SK hynix puts KV cache tiering on a hardware roadmap

Semiconductor Week 38: memory crosses half the market, and SK hynix puts KV cache tiering on a hardware roadmap

Source: The Semiconductor Newsletter, Week 38 2026, via starred Gmail · Post Raw: raw/gmail/2026-09-23-starred.md

TL;DR

The week's single most decision-relevant number for anyone tracking inference cost: global semiconductor revenue hit $425 billion in Q2 2026, and memory crossed 50% of the market for the first time. The industry's centre of gravity has moved from logic to memory, which is the arithmetic consequence of a workload whose bottleneck is bytes moved rather than FLOPs executed. Alongside it, SK hynix announced an AI memory architecture spanning HBF (high-bandwidth flash), PIM (processing-in-memory), and explicitly "tiered KV cache management." A memory vendor is now naming the KV cache as a product tier. Also this week: NVIDIA's Vera Rubin NVL72 delivered up to 3.7x GB300 throughput in an MLPerf preview, Micron demonstrated a 512 GB DDR5 RDIMM at 9,200 MT/s, CXMT moved an 11.95 nm G5 DRAM platform into mass production, and OpenAI took its Jalapeño accelerator from RTL to tape-out in nine months using LLMs.

The items that matter to this wiki, in order

1. Memory exceeds 50% of a $425 billion quarter. For most of the industry's history, memory was the commodity half and logic was where the margin lived. This wiki's compute-economics page has been tracking the inversion through its symptoms: SemiAnalysis's 4-hi HBM study (09-14) showing vendors choosing bandwidth over capacity, the Engram offloading study (09-18) reporting NVIDIA despec'ing Rubin Ultra from 1024 GB to roughly 200 GB of HBM per chip, and the 09-22 four-regime inference decomposition showing that decode attention is context-bound and decode experts are weight-bound, neither compute-bound. The revenue split is the market pricing that fact.

2. SK hynix names tiered KV cache management as a memory-architecture feature. This is the item to carry forward. Everything on the KV cache page is a software argument about how to spend a memory budget: compress it, evict it, quantize it, share it across depth, offload it to a cheaper tier. KVMEM (09-23) built a GPU-to-RAM-to-NVMe paging hierarchy for agent memory entirely in software, on a laptop. SK hynix pairing HBF (flash with high-bandwidth-memory-style interfaces, aimed at capacity behind HBM) with PIM (compute placed inside the memory array, so some operations never cross the bus) and an explicit KV tier means the hierarchy those software systems improvise is becoming a hardware product line. If the tiering moves into the memory stack, every software offloading result on the KV page acquires a hardware baseline it currently lacks, and some of them become redundant.

3. Vera Rubin NVL72 at up to 3.7x GB300 throughput in MLPerf preview. Vendor-run preview numbers, so discount accordingly, but the direction is consistent with the despec story: the gain is coming from the rack and interconnect rather than from per-chip HBM capacity, which was cut. A 3.7x throughput improvement on a chip with roughly a fifth of the previously-planned HBM is only coherent if the workload is being reorganised around smaller resident caches, which is precisely the four-regime disaggregation SemiAnalysis described on 09-22.

4. OpenAI's Jalapeño accelerator went RTL to tape-out in nine months, with LLMs completing the RTL. Two readings and both are worth holding. The capability reading is that hardware design is now a domain where the agentic loop produces a physical artifact, which is a stronger claim than a benchmark score. The caution is that "LLMs completed the RTL" is a vendor framing with no defect-rate or verification-coverage number attached, and RTL is a domain where a silent error is expensive in a way a hallucinated paragraph is not. It belongs next to AIDE² (09-23) as the second self-improvement-adjacent result this week where the loop's output is not text.

5. The supply-chain build-out continues on two fronts. India drew up to $12 billion in new semiconductor supply-chain proposals, with Tata Electronics forming partnerships with Fujifilm (materials) and Besi (advanced packaging). Anderon finalized a $1 billion CHIPS award for a 300 mm quantum wafer foundry. China-side, CXMT's 11.95 nm G5 DRAM entering mass production and IMECAS demonstrating a DUV process path for stacked-nanosheet gate-all-around devices are the two items that matter for the export-control thread, because both are capability advances that route around EUV.

6. Funding and consolidation. Euclyd raised more than €200 million for energy-efficient AI infrastructure. Winbond acquired Infineon's NOR flash and F-RAM business for $1.12 billion. Air Products committed $250 million to Arizona semiconductor gas infrastructure. Photonic proposed a CAD 500 million multi-tenant facility in Canada. SK hynix established a Silicon Valley venture brand to invest in AI infrastructure, which is a memory vendor buying optionality on the software layer that consumes its product.

7. The constraint everyone named. McKinsey identified custom silicon, power, and talent as the core AI-infrastructure constraints, and a separate item in the same issue reports parallel talent gaps in both data centres and fabs. An AI Energy Management Alliance launched targeting grid-responsive data centres. Power as a first-order constraint is now appearing in the same paragraph as compute in industry analysis rather than as a footnote.

What this changes for the wiki

The KV cache page needs a hardware section it does not have. Every entry on it assumes the memory hierarchy is fixed and the question is how to fit inside it. SK hynix's announcement, read against the Rubin Ultra despec and the Micron 512 GB RDIMM, says the hierarchy itself is the variable now. The open question to carry: if PIM can execute attention-adjacent operations inside the memory array, which of the cache-compression techniques on that page are solving a problem that moves?

And the memory-majority revenue split is the cleanest available answer to why routing and compression results keep converging on data movement. The 09-22 routing-tax finding, that an Opus to Sonnet to Opus route costs 6.19 against 4.15 for never switching because each hand-off reprocesses context, is a data-movement result wearing a routing costume. So is Flash-dLLM (09-23), which found GPU memory I/O rather than arithmetic to be the dominant bottleneck once caching and parallel decoding are combined. When memory is over half the industry by revenue, the surprise would be if the software bottlenecks were anywhere else.

Related pages