DeepSeek has rewritten its core training operators for Huawei's Ascend architecture and released them for free, directly targeting the most fortified segment of NVIDIA's CUDA ecosystem.
A lab training its own large models rewrote its core operators to run on a rival chip's instruction set and released them for free. On September 30, DeepSeek announced the full open-sourcing of infrastructure components for Huawei's Ascend platform, with Huawei providing engineering support throughout the development process.
First, clarify what is being migrated. This release includes six components, each corresponding to a previously open-sourced component for the NVIDIA platform: TileLang for Ascend, the matrix operation library DeepGEMM-Ascend (supporting BF16, FP8, and FP4 precision), the distributed communication library DeepEP-Ascend for large-scale expert parallel models, the standard vector operator library TileKernels, the long-context attention operator FlashMLA, and the data filtering component DeepSelect. The most critical commitment is that most operators used in training the DeepSeek V4 series models now have high-performance Ascend versions.
Why TileLang is the primary focus
Among the six components, DeepSeek specifically designates TileLang as the core. Developed by a team from Peking University's School of Computer Science and open-sourced in January 2025, this language is positioned as a programming bridge between large model algorithms and underlying chip hardware. DeepSeek cites three advantages: compared to CUDA, it offers simpler programming, shorter code logic, and higher development efficiency; compared to other high-level parallel languages, its programming model fully utilizes chip features to reach hardware performance limits. Officially, TileLang now underpins the implementation of most operators in DeepSeek V4 series model training—simultaneously satisfying the two hardest conditions in ecosystem migration: ease of writing and sufficient speed.
Why the language layer matters more than the operators: CUDA's moat is built on a developer base in the millions and two decades of tooling. Porting operators one by one will never keep pace. A language that lets Python-fluent engineers write kernels near hardware limits changes the supply of talent: developers no longer need a decade of GPU-specific training. This steepens the growth curve for the Ascend ecosystem. This is the first time a frontier lab has backed such a language with its own primary training workloads.
Tuning a 128-chip Ascend 950 supernode
Beyond open-sourcing, DeepSeek and Huawei jointly tuned a supernode configuration comprising 128 Ascend 950 chips. Optimization targeted both ends: single-chip compute utilization and inter-chip data transfer efficiency. In large-scale model training, these two factors often conflict: maximizing compute can bottleneck communication, while minimizing communication can leave compute idle. Officials state that in key scenarios, compute and communication performance have approached the theoretical limits of Ascend hardware. The timing is notable: Huawei released its next-generation Ascend chips two weeks ago, and model vendors immediately followed with rewrites. For the first time, the cadence of hardware and software aligned within the same quarter.
Zooming out on the timeline clarifies the motivation behind this collaboration. Multiple media reports indicate that the new-generation Ascend chips are in short supply. Huawei plans to limit overseas sales to prioritize domestic customers. Huawei Rotating Chairman Xu Zhijun has previously stated plainly that the company cannot accept a future where it does not control its own destiny. Pairing supply-constrained chips with a rapidly maturing software stack gives the Ascend ecosystem, for the first time, both assured capacity and usable tools—two conditions that have never coincided before.
6
Open-source components
3
DeepGEMM precision formats
128
Supernode chip scale
The three figures above represent: the number of Ascend platform components open-sourced in this release, the number of compute precision formats supported by the matrix operation library DeepGEMM-Ascend, and the number of Ascend 950 chips used in the jointly tuned supernode configuration.
What is still missing for true replacement
Before claiming full replacement, the dependency list must be examined. DeepGEMM-Ascend still runs on Huawei's proprietary CANN compute architecture, and TileKernels requires either CUDA or CANN as the underlying layer. Developers migrating are switching from one vendor's stack to another, not escaping vendor lock-in. Additionally, the Ascend 950 lacks large-scale independent third-party benchmarks, and no one has quantified how much migration effort external communities can save using this toolset.
However, one signal has already materialized: industry research firm SemiAnalysis notes that CANN is the only compute stack, aside from CUDA, capable of supporting DeepSeek V4 on its release day. DeepSeek has previously prioritized early access to V4 for domestic chipmakers. Model vendors treat 'whose chips run the new model on day one' as a strategic statement, which reveals the depth of collaboration more than open-source code alone. Software adaptation has shifted from a catch-up exercise for laggards to a default item in the release cadence.
My assessment is that the true test for this open-source suite lies in DeepSeek's next-generation model. Multiple financial media outlets have reported that a data center in Inner Mongolia is planned with at least 160,000 Ascend accelerators. If the primary training workload for the next generation truly shifts to Ascend, the narrative of CUDA replacement upgrades from 'components are ready' to 'the closed loop works.' The falsification condition is equally clear: if next-generation model training remains primarily on NVIDIA platforms, with Ascend handling only inference or edge workloads, this six-component suite remains at the level of an ecosystem demonstration.
For engineering teams uncertain about whether to trial the platform, this open-sourcing effectively includes decision-making materials: all six components are hosted in public repositories, with source code, build instructions, and benchmark scripts directly downloadable. To evaluate migration costs, teams do not need to wait for commercial negotiations; they can run key operators on a small model in a single-node Ascend environment to obtain their own measured figures on performance gaps and engineering effort. Ecosystem debates yield no results, but empirical testing does.
A language that is easy to write and fast enough, combined with an operator library stress-tested by real training workloads, is closer to the tangible form of ecosystem building than any number of 'self-reliance and controllability' slogans. The next move to watch is not the launch event, but the training logs of the next-generation model.