Chinese AI lab DeepSeek has partnered with Huawei to release a suite of open-source programming tools for Huawei's Ascend AI chips, a move that takes direct aim at the software ecosystem that has underpinned Nvidia's dominance in artificial intelligence computing.

The announcement was made on DeepSeek's official WeChat channel and confirmed by Reuters, Bloomberg and The New York Times. According to The Decoder, the release includes libraries for computation and for moving data between chips — and all of it is open source. DeepSeek said Huawei "fully supported" the work.

For readers tracking the breaking AI news out of China's chip industry, this is one of the most consequential software releases of the year: it is the clearest sign yet that China's leading AI lab is building its stack around domestic silicon rather than Nvidia GPUs.

What DeepSeek Actually Released

At the center of the release is TileLang, an open-source programming language for AI chips that was originally developed by researchers at Peking University. DeepSeek has been using the language internally for about a year, according to The Decoder, and first tested it on older Nvidia chips before bringing it to Huawei hardware.

DeepSeek's argument is that anyone trying to build an independent software ecosystem for AI chips needs a universal language that is easy to program while still extracting full performance from the hardware. In the company's view, TileLang offers a simpler programming model than CUDA, Nvidia's proprietary platform that has been the de facto standard for AI development for nearly two decades.

The New York Times reports that TileLang is now the main tool in DeepSeek's own work toward artificial general intelligence — a notable endorsement, given that DeepSeek's models are trained and served at frontier scale.

Beyond the language itself

Alongside TileLang, DeepSeek shipped open-source libraries covering computation and inter-chip data movement, the two workloads that dominate large-scale AI training and inference. The company also said it jointly optimized a "supernode" with Huawei — a cluster of 128 Ascend 950 chips designed to work as a single system.

Why This Is Aimed Straight at Nvidia

Nvidia's advantage has never been chip design alone. The company's real moat is CUDA and the ecosystem around it: an estimated four million developers worldwide build with CUDA, and that installed base is something rivals like AMD have struggled to cross even when their hardware looked competitive on paper, The Decoder notes.

That is what makes the DeepSeek-Huawei release different from previous Chinese chip announcements. Rather than only building faster hardware, China's AI industry is now trying to replicate the software layer that locks developers in.

The New York Times reports that Chinese model makers such as Z.ai and Moonshot AI have, until now, moved faster than the country's chipmakers — meaning China's best models were still largely dependent on foreign compute. Huawei wants to close that gap. Two weeks before DeepSeek's announcement, Huawei unveiled a new generation of AI processors and supernode systems and said they would be widely used for model training next year.

The export-control backdrop

The partnership cannot be separated from US export controls, which have progressively cut Chinese companies off from Nvidia's most advanced GPUs. Huawei's rotating chairman Eric Xu has said the company cannot accept a future that hinges on whether others are willing to sell chips to China, according to The Decoder. DeepSeek's move to publish its chip software openly also makes it harder for any single vendor — Chinese or American — to control the layer everyone depends on.

Is the CUDA Moat Actually Cracking?

Interesting context comes from research firm SemiAnalysis, which recently tested Jalapeño, OpenAI's in-house inference chip, and concluded that the CUDA moat is "potentially dead" — noting that OpenAI gets new models running on its own hardware remarkably quickly, and that Jalapeño beat Nvidia's Blackwell on performance per watt in most of the scenarios tested, as The Decoder reports.

But the analysts added important caveats. They only tested relatively easy-to-optimize scenarios with roughly 8,000 input tokens and 1,000 output tokens, and they have not yet run AgentX, a benchmark measuring how AI agents handle multistep tasks — exactly where SemiAnalysis found Nvidia well ahead as recently as August.

What it means for the industry

The realistic read is not that CUDA collapses overnight. Four million developers, two decades of libraries and countless production deployments do not migrate quickly. But the combination of DeepSeek's software, Huawei's new chips, and custom silicon efforts like OpenAI's suggests the software lock-in that made Nvidia the default is being chipped away from multiple directions at once.

For China, the payoff is strategic: a viable, open alternative software stack makes domestic chips usable at scale, reducing the leverage of US export policy. For everyone else, an open-source TileLang ecosystem could eventually lower the cost of targeting non-Nvidia hardware anywhere in the world.

DeepSeek and Huawei have not disclosed performance benchmarks comparing the new stack against CUDA on equivalent hardware, so concrete speed and efficiency claims remain unproven. What is clear is the direction: the most-watched AI lab in China is no longer treating Nvidia's software as an unavoidable dependency.

Stay Ahead of AI

Missed a story? Read more AI news →