DeepSeek and Huawei Expand TileLang for Ascend AI Chips
DeepSeek and Huawei expand TileLang as the two companies develop open-source programming infrastructure for Huawei’s Ascend AI chips. DeepSeek disclosed the partnership on September 30, naming the high-level language alongside related compute and communication libraries intended to make Ascend hardware easier to program.
The work targets one of Nvidia’s strongest competitive advantages: the software ecosystem surrounding its chips. Huawei has built capable AI processors, but developers also need compilers, libraries, debuggers and optimized kernels before Ascend can serve as a practical alternative for demanding model workloads.
The expanded stack brings together four components:
- TileLang’s Python-style kernel language
- Ascend-specific compiler backends
- Compute and communication libraries
- Open tooling for performance optimization
DeepSeek and Huawei Expand TileLang for Ascend Chips
Reuters reported that DeepSeek announced the partnership through its official WeChat account. The company said it was open-sourcing infrastructure for Ascend, including TileLang and supporting libraries, as Chinese developers seek a more independent software foundation for AI computing.
TileLang itself predates the announcement, and its Ascend adapter has been open source since September 2025. The new disclosure matters because it connects DeepSeek’s model-engineering needs with Huawei’s hardware roadmap and places additional shared infrastructure around an existing developer project.
The repository already includes kernels for DeepSeek V4 workloads, showing that the collaboration extends beyond a strategic promise. DeepSeek has increasingly optimized its models for Ascend systems, while Huawei has opened more of the programming layers that sit between model code and its processors.
How TileLang Maps AI Kernels to Ascend
TileLang is a domain-specific language for writing the small, performance-critical programs known as kernels. These kernels handle operations such as matrix multiplication, attention, normalization, sorting and activation functions that consume much of the compute used during model training and inference.
The Ascend version uses Python-style syntax on top of TVM compiler infrastructure. Its two technical paths can generate code through Ascend C and Huawei’s PTO instruction architecture or through AscendNPU IR, giving developers different routes from a high-level kernel description to executable NPU instructions.
The project supports common operations including GEMM, Flash Attention, sparse attention, convolution and cross-entropy loss. It has been tested on Huawei’s A2 and A3 processors, and it integrates with PyTorch through an Ascend-specific bridge for developers working in established machine-learning frameworks.
Huawei describes PTO as a tile-based instruction architecture with more than 120 virtual instructions across eight categories. Tile-level abstraction is intended to preserve operator source code across hardware generations while still allowing manual control when developers need to tune memory movement or computation.
The Software Layer Challenges Nvidia CUDA
Nvidia’s CUDA advantage is not limited to a programming language. It includes mature libraries, profilers, documentation, optimized kernels and years of developer experience. A rival chip can offer competitive specifications and still struggle if engineers must rewrite software or sacrifice reliability to use it.
DeepSeek says TileLang offers a simpler programming model than CUDA and reduces code complexity. The relevant test will be whether that abstraction consistently reaches high utilization on real Ascend clusters without forcing developers to drop into hardware-specific code for every important workload.
Communication libraries are especially important for large models distributed across many processors. Training and serving can stall when devices wait for data, so the partnership must optimize both individual kernels and the movement of tensors between processors, servers and large Ascend supernodes.
Huawei’s broader platform already supports PyTorch, Triton, vLLM and other third-party projects. The company says more than 40 models have been natively pre-trained on Ascend and that external contributors now represent most CANN developers, evidence that it is trying to turn a proprietary stack into a wider ecosystem.
Open Source Will Decide Ascend Adoption
Publishing code lowers the barrier for researchers and infrastructure teams to inspect, modify and benchmark the stack. It also lets developers identify missing operators and compatibility problems before committing large budgets to Ascend clusters, a process that closed tools make harder.
Open repositories do not guarantee an open ecosystem. Ascend deployments still depend on Huawei’s CANN toolkit, device availability, documentation and access to suitable compute. The health of issue trackers, release cadence and independent contributions will matter as much as the initial code release.
The partnership could give DeepSeek greater influence over the software used to run its models on domestic hardware. Huawei, meanwhile, gains a prominent model developer that can expose bottlenecks and supply demanding workloads for compiler and kernel improvements.
If the tools mature, developers may gain a credible route to deploy large models without centering every optimization on Nvidia hardware. The near-term result is narrower but still consequential: DeepSeek and Huawei are treating the programming stack, not just chip design, as the central battleground for AI infrastructure.
Background Reading