Nvidia’s real moat was never the silicon. It was CUDA — a programming platform with an estimated four million developers, so deeply embedded in the machine-learning stack that “porting off CUDA” used to be a polite way of saying “good luck, that’s a decade of work.” On Wednesday, DeepSeek handed Huawei a free, open-source shortcut out of that moat, and the timing is almost too neat.
Six modules, one WeChat post
The release landed on a Wednesday, announced in a post on DeepSeek’s WeChat account and confirmed by Reuters. The headline: six open-source software modules for Huawei’s Ascend line of AI chips, built with “full support” from Huawei, and free to download.
The three that matter most:
- DeepGEMM-Ascend — handles the matrix multiplication and other heavy lifting that DeepSeek’s models are built on. It supports BF16, FP8 and FP4 operations, and — the detail that will actually keep engineers up at night — it uses the same programming interfaces as DeepSeek’s existing DeepGEMM library. You keep your familiar API. You just point it at a Chinese accelerator.
- DeepEP-Ascend — the communication layer for training and inference, including routing data to the different “experts” in a mixture-of-experts model and stitching their outputs back together. This is the part that decides whether a cluster of chips actually behaves like one machine.
- TileLang — the high-level programming language DeepSeek has been pushing as a simpler alternative to CUDA. The September 30 update adds native support for the Ascend 950, including code generation, automatic scheduling and synchronisation.
The two companies also co-developed and tested the stack on a “supernode” system built from 128 Ascend 950 chips — the unit of scale that decides whether you can train a frontier model on domestic hardware at all.
Why this is more than a chip story
Here’s the thing most coverage gets wrong by framing it as “China builds another AI chip.” The chip isn’t the point. Huawei’s latest accelerators still trail Nvidia and AMD on paper — everyone who’s looked at the spec sheets says so. The point is the software.
CUDA’s advantage was always that it was mature, well-documented, and four million developers deep. You didn’t pick CUDA because it was the fastest; you picked it because everything was written for it. So the most effective attack on the moat isn’t a faster chip — it’s a simpler language that lets you write the same work in less code, and an open-source library set that means the ecosystem can start accumulating on a second platform without a single line of paywalled, vendor-locked tooling.
DeepSeek describing TileLang as offering “a simpler programming model” than CUDA is doing a lot of quiet work in that sentence. They’re not promising to match CUDA feature-for-feature. They’re promising that a competent engineer can get a working, optimised kernel on an Ascend 950 without first doing a PhD in GPU scheduling. And they built it on top of Huawei’s existing CANN platform, so there’s real infrastructure underneath, not a sandbox demo.
The numbers that make it real
The software release lands on top of some genuinely substantial hardware commitments, which is what separates this from a demo:
- Huawei is currently deploying a 256,000-card Atlas 950 SuperCluster, and its newer architecture is designed to scale to as many as one million NPUs.
- DeepSeek plans to deploy at least 160,000 Huawei accelerators in a data centre in Inner Mongolia.
- Huawei expects its AI systems to be widely used for model training by 2027.
- The two companies had already collaborated on DeepSeek’s V4 model, released in preview in April, with Ascend support — and Huawei says its chips were used for part of training the lighter V4-Flash variant.
And the financials: DeepSeek’s annualised revenue run rate reportedly hit $1 billion last week, and the company is raising at a valuation approaching 500 billion yuan (about $74 billion). This is not a hobby project.
The market-share bomb that came with it
The software news broke the same week Huawei’s rotating chairman Eric Xu made a claim that deserves a moment of stunned silence: that Ascend has surpassed Nvidia in Chinese market share.
“It’s pretty hard to collect data about the market share of Nvidia in China, but based on the data we have collected, Ascend has surpassed Nvidia,” he said. His reasoning is less about specs than about something colder: “Even though our chips may be less advanced, at least their supply is assured, so that you don’t have to worry about chip supply day in and day out.”
That’s the actual driver. The US export-control whiplash — licensing requirements, then reversals, then a grudging H200 carve-out — taught Beijing that foreign silicon can be switched off at the stroke of a pen. So the Chinese government has been pressing datacentres to move to domestic accelerators, and “we’re a little slower but we won’t get banned” is, for a state that’s been burned, a genuinely compelling sales pitch.
My take
I find the timing almost suspicious, in the good way. Two weeks after Huawei unveiled its next-gen Ascend chips, its chairman claims the domestic chip has taken the lead in China — and then DeepSeek drops an open-source toolkit that makes those chips programmable by the same developers who’d been stuck on CUDA. The chip and the software are two halves of one move, and the second half is the one that actually changes the industry.
The honest caveat: Huawei itself says it doesn’t have the capacity to meet Chinese demand, let alone export. This is a China story first. But the principle — that the moat is the toolchain, not the transistor, and that an open-source, simpler-language strategy can chip away at a four-million-developer ecosystem — is one the rest of the world should be taking seriously. CUDA has been the default for so long that “the CUDA moat is still safe” stopped being an argument and became a habit.
DeepSeek, the lab that kept embarrassing the frontier labs with models that cost a fraction to train, has just made the single most boring-sounding strategic move of the year: open-source the plumbing. And it might be the one that matters most.
Sources: Tom’s Hardware / Reuters, The Next Web, The Register.
