The DeepSeek Kernel Engineer's Essay: A Practitioner's Timeline for Agent-Written CUDA
Source: Shengyu Liu (刘胜与 / interestingLSY), kernel engineer at DeepSeek · surfaced via the X home feed, amplified by @teortaxesTex, @Thom_Wolf, @AndrewCurran_, @hsu_steve, @MaxForAI, @CrazyShyyt · Date: 2026-09-15 Raw: X home feed
TL;DR
The most-shared item on the X feed today, by a wide margin and across language communities, is a long personal essay by a DeepSeek kernel engineer whose work is in DeepSeek V4.1's main attention kernel. This wiki normally ignores social commentary. This one is recorded because it is a first-person capability report from inside the exact cluster gpu-kernels has been assembling all year, and because the person making it has every professional incentive to say the opposite.
The capability claim, with a timeline. A year ago, AI could help him look up documentation, read code and hunt bugs. Now it reads CUDA, PTX and SASS directly, uses professional profiling tools to analyse stalls per instruction, and iterates on the optimization. His estimate: six to twelve months until AI-written kernels match or surpass his own. PTX and SASS matter here. PTX is NVIDIA's intermediate assembly and SASS is the actual machine code the hardware runs, and reasoning at that level is where the last 20% of kernel performance lives. It is also the level with the least public documentation, which is precisely why it was assumed safe.
The labour claim is more specific than "AI takes jobs." He does not expect unemployment. He expects to stay in the industry with the work transformed: from writing kernels to piloting an agent that writes kernels. Industry demand, he says, has already drifted from "people who can write high-performance kernels" to "people who can use AI to produce high-performance kernels faster." The loss he names is the craft itself. He used to spend an afternoon tuning performance like a speedrunner chasing his own record, and his phrase for what comes next is leaving his talent behind in yesterday to become an agent's driver.
The second half is a distribution argument and it is why the essay travelled. He argues the endpoint is one of two extremes. Either frontier AI becomes cheap infrastructure everyone can access, or a handful of companies own the strongest models and everyone else runs a weakened version. He describes the lock-out as a dead loop: you need the strongest AI to earn the resources, and you need the resources to access the strongest AI. He states plainly that he does not believe Anthropic or OpenAI will voluntarily open their frontier, and that this is why he stays at DeepSeek: to research stronger, faster, more affordable AI and open-source it, pulling the world at least slightly away from the second outcome. He is accelerating his own obsolescence deliberately, on the grounds that the alternative is worse.
How this relates to the rest of the wiki
This is practitioner testimony for a research cluster that had none, and it arrived the same day as that cluster's sixth paper. gpu-kernels holds five results on agentic kernel optimization: AccelOpt (04-20) took Trainium peak-throughput utilization from 49% to 61% while matching Claude Sonnet 4 at 26x lower cost, JAXBench (08-03) showed curated TPU documentation moved Gemini 3 Flash from 5.8% to 37.3% per-sample correctness, and MaxKernel (09-07) built three autonomy operating points over one shared sub-agent pool and matched expert hand-tuned baselines on 50 TPU kernel tasks. All three are benchmark results produced by people trying to show that agents can write kernels. Today's essay is the same claim from someone trying to describe his own displacement, and today's Dream-RSI evaluates on GPU kernel engineering as one of three domains while optimizing the search strategy rather than the kernel. Six results, one practitioner report, and the practitioner's timeline is shorter than the papers'.
The MaxKernel connection is exact and worth stating. That paper's three operating points were a human-in-the-loop design agent, a fully autonomous metric-and-trace-driven optimization loop, and a graph-based autonomous search. Liu is describing the middle one arriving in production: an agent that reads a profile trace, finds the stall, and iterates. MaxKernel's contribution was varying autonomy as a design parameter; Liu's observation is that the industry is choosing the high-autonomy point for economic reasons before the research has settled which point is best. The paper asked which operating point wins and the labour market is answering it first.
It also supplies the missing character in this week's pacing the frontier argument. Dario Amodei's essay claims AI is accelerating "driven primarily by AI's growing ability to build the next generation of AI," and the debate that followed was conducted entirely between executives, safety researchers and markets. Liu is the person the claim is actually about: the AI is getting better at building the next generation of AI by getting better at the attention kernel he wrote. His answer to the acceleration is not to slow down but to open-source, on an explicitly distributional argument that the pacing debate does not engage with at all. The safety framing asks how fast; his framing asks who owns it, and those are different axes that this week's discourse has collapsed.
Caveat on how this reached the wiki. The essay was read through six secondhand summaries rather than in the original, including two in Chinese and several from accounts that routinely overstate their sources. The claims recorded above are consistent across all six, including the independent Chinese-language accounts, which is the main reason to trust the substance. The six-to-twelve-month timeline is one engineer's estimate about his own field, not a measurement, and it should be held as such. The right falsifier is public: an open kernel benchmark on which an agent beats a named expert-tuned baseline on NVIDIA hardware at the SASS level.