DeepSeek and Huawei just gave developers a reason to write AI code without touching Nvidia's CUDA.
The two companies released open-source programming libraries for Huawei's Ascend AI chips on September 30. The toolkit includes DeepGEMM-Ascend for matrix multiplication, which supports BF16, FP8, and FP4 and mirrors the APIs of DeepSeek's existing DeepGEMM library, plus DeepEP-Ascend for the chip-to-chip communication that keeps mixture-of-experts models fed with data. DeepSeek also added native Ascend 950 support to TileLang, the high-level language it's pitching as a simpler alternative to CUDA for writing optimized kernels. The work builds on Huawei's CANN software stack and was tested on a supernode system built from 128 Ascend 950 chips.
Nvidia's real advantage was never just silicon - it's years of mature CUDA tooling that makes GPUs easy to program, and that ecosystem has been central to its dominance in AI computing. Open-sourcing these libraries chips away at that advantage, giving developers a documented path to full Ascend performance instead of reverse-engineering undocumented behavior.
TileLang still runs on Nvidia GPUs too, so this reads less like a wall going up around Huawei's chips and more like a bet that good tools eventually win users over regardless of whose logo is on the silicon.