DeepSeek released a batch of things over the last two days: six repositories plus one technical report in Chinese, all around Huawei's Ascend. I cloned all six, read them, and tried them on my own machine.
This post covers three things: what was actually released, whether those impressive-looking numbers hold up, and exactly where my machine gave up.
The short version
- Six components were open sourced this time, and every one of them has a CUDA counterpart: a language for writing kernels, matrix multiply, multi-card communication, a general kernel library, an attention kernel, and a TopK kernel. Together that's filling in a column, not shipping a new model.
- The numbers are strong: dense matrix multiply at 99.8% of the hardware limit, sparse attention prefill at 95%, multi-card communication at 90–95% of physical bandwidth. But every one of these was measured on PoC hardware Huawei provided to DeepSeek. The commercial version is planned for mid-October. No third party can reproduce them today.
- The most valuable thing in this release isn't the announcement — it's the Chinese technical report published alongside it. It lays out the trade-offs of writing a sparse attention kernel on Ascend step by step: why one AI Core handles 64 heads, why you must pair each CUBE with two vector cores, and why one layout conversion has to detour through UB. I've copied the parts I found most revealing below.
- What I could verify locally: all six repos cloned, I read and counted the code, and I could not run a single kernel. The build dies at line 7 of
setup.py. That section is below too — it's measured, not paraphrased.
(Component-by-component comparison. The rightmost column is the size of the Ascend-specific part, which I counted from the source.)
1. Getting this layer straight first
Between a large model and the silicon sits a software layer: it tells the chip how to compute (kernels) and how to move data between cards (communication libraries). NVIDIA's version of that layer is CUDA; Huawei's is CANN. Anyone can write the model on top, but this layer has to be rewritten in the chip vendor's own language.
When DeepSeek V4 launched in April, Huawei got there first with Ascend support — but that delivery was about making it run. The pieces that matter most for performance were still open sourced only in their CUDA form. One media analysis at the time put it this way: moving from CUDA to CANN involves rewriting kernels at scale, restructuring the training pipeline, and re-tuning accuracy and performance — not less work than training the model itself. (Source: 21st Century Business Herald, 2026-04-24, in Chinese.)
That is exactly what this release is. Which is why the sentence in the announcement — "all components correspond one-to-one with the previously open-sourced NVIDIA-platform components" — is the most substantive line in it.
2. The six components, one by one
TileLang's Ascend 950 backend. TileLang is a language for writing high-performance kernels in Python, led by a team at Peking University and used for most of the kernels in DeepSeek V4's training. This time it landed in the main repository, not as a bolt-on adapter: PR #3308, "Introduce Ascend 950 backend", merged on 30 September and shipped with v0.1.15. The official support matrix lists Ascend 950 as Supported, built from source with USE_ASCEND=ON USE_CUDA=OFF, requiring CANN, the bisheng compiler and torch_npu on the machine, with torch.npu.is_available() returning True.
DeepGEMM-Ascend (matrix multiply). Same package name and same API as the NVIDIA version, so existing code keeps working. The repo was created on 29 September.
DeepEP-Ascend (multi-card communication). MoE dispatch and combine — sending tokens to the right expert and collecting them back. The kernels are all header-only .hpp files sitting on Huawei's HCCL/HCOMM, UBMEM and URMA. It's also the only Ascend repo in this batch that was built from scratch.
TileKernels v2.0.0 (general kernel library). The small but frequently used kernels: MoE routing, quantization, Engram, hyper-connections. The README's own words: "added an Ascend backend, selected automatically at runtime; the same Python API runs on both NVIDIA GPUs and Huawei NPUs."
FlashMLA (attention). Sparse attention prefill and decode.
DeepSelect (TopK). The step that picks compressed blocks in sparse attention, plus the TopK the sampler needs. The authors claim 2–20× faster than native torch.topk.
3. The numbers are strong — read them together with their sources
(Performance figures and where they come from. Bar length = the fraction of that hardware's theoretical limit; further right means closer to the ceiling.)
The key readings:
- Dense matrix multiply at 99.8% of the hardware limit (BF16, FP8 and FP4 all land around there). FP4 measured 1701 TFLOPS on a 4096×7168×16384 shape against a 1730 TFLOPS hardware limit, on Ascend 950DT with CANN 9.2.0.
- Sparse attention prefill at 410 TFLOPS, 95% of the theoretical limit; decode at 360 TFLOPS, 83%.
- MoE communication: at EP8, dispatch 373–375 GB/s and combine 345–347 GB/s; within 32 cards it reaches 90–95% of physical bandwidth. At EP128 dispatch drops to 313–320 and combine to 272–278, with the vendor saying combine is still being optimized.
The sourcing point has to be stated plainly: DeepEP-Ascend's own README says these numbers were measured on a PoC HDK that Huawei provided to DeepSeek, with additional manual configuration that is not part of any public release. Huawei's commercial HDK for Atlas 850E is planned for mid-October. A third party cannot get the same hardware and configuration right now, so this set of numbers should be read as a vendor claim, not as a reproducible public benchmark.
Two pieces of hardware background, both from official statements: the Atlas 950 supernode supports 8192 950DT chips across 160 cabinets (128 compute, 32 interconnect), with an FP8 rating of 8 EFLOPS and FP4 up to 16 EFLOPS, and interconnect bandwidth of 16 PB/s — Huawei's own line being that this single product's total interconnect bandwidth is more than ten times the peak bandwidth of today's entire global internet. Availability is Q4 2026. A single Ascend 950DT is planned at 144 GB of memory, 4 TB/s of memory bandwidth and 2 TB/s of interconnect bandwidth.
That is the basis for the fine print on DeepSeek's side: the pricing cadence of V4-Pro is tied to the shipping schedule of the Ascend 950 supernode.
Further out, there are moves in progress: Bloomberg reported on 4 September (relayed by Guancha) that DeepSeek plans to deploy at least 160,000 Ascend 950DT chips for inference in an Inner Mongolia data center, with a gigawatt-scale power plan and full delivery possibly taking more than a year. Huawei announced the Ascend 960 supernode at its 17 September Connect conference — 4096 cards, replacing 48,000 800G optical modules with 5,500 in-house optical engines. Both of these are third-party reporting or vendor announcements; I have no independent verification to offer.
4. That Chinese technical report is the part worth reading
FlashMLA shipped with a paper called 《Ascend Sparse Attention Forward 算法与优化技术简析》("A brief analysis of the Ascend sparse attention forward algorithm and optimization"), in Chinese, released today and sitting in the repo. It lays out line by line why the kernel is written the way it is. A few of the parts that stuck with me:
One AI Core handling 64 query heads is not arbitrary. The report puts the compute-to-memory ratio on the table: with CUBE saturated (4096 multiply-adds per cycle), L2 total bandwidth of 5 TB/s, 32 cores and 1.65 GHz, the head count has to be 64 or more before compute becomes the bottleneck. At the same time, 128 heads would overflow L0C (a 128×512 float32 output is 256 KB), forcing the accumulator onto the vector core, which is extremely hard to write; empirically M=64 is already enough to saturate CUBE.
Every CUBE needs two vector cores; one is not enough. All three reasons are hard constraints: a single vector core can have at most 16 MTE2 copy requests in flight; decode has to read FP8 or FP4 KV and dequantize to bf16, and one vector core can't dequantize fast enough to keep up with CUBE; and FixPipe writes out at only 128 bytes per cycle per core, which by the formula forces B_TOPK above 73 — in practice they use 96 or 128. Taken together, the second vector core is not redundancy.
A purely arithmetic optimization saves a large round of data movement. The conventional online softmax approach may need to rescale the accumulated output against a new maximum every iteration. Their approach sets a threshold of 6 for "does this need rescaling" and skips the rescale when it isn't triggered — a trick they call skip-scale. The payoff is bigger on an NPU than a GPU, because on an NPU scale-O has to touch three units: FixPipe, UB and SIMD. The hard part is that the signal is produced by the vector core while the consumers are the CUBE and the vector core, which they solved with an SS buffer plus a cross-core flag.
The KV layout conversion goes the long way around because of a memory access pitfall. CUBE wants NZ layout while KV sits in memory as ND. The normal approach is to convert in flight during the copy, but they fetch KV sparsely token by token, which rules that out. They tried copying tokens as 32-byte blocks directly, and found that the BIU inside the AI Core only coalesces accesses when both source and destination are contiguous — "source contiguous only" doesn't coalesce, L2 bandwidth utilization drops, and the bottleneck becomes L2 bandwidth itself. In the end the path is: memory to UB, SIMD vector instructions to convert into NZ, then into L1.
The report also covers gather2 (fusing two token copies into one to work around the MTE2 queue depth of 16), L1 dual-bank balancing (splitting each tensor into two, one per bank), and the fact that icache is only 32K (CUBE), 16K (VECTOR) and 8K (SIMD VF) — which is why they use hardware loop instructions instead of unrolled loops. None of these details appears in the announcement.
One companion piece worth reading: Huawei's Ascend community has a long tutorial on TileLang from beginner to expert (published 2026-05-08) that lists the expert-mode conversion as six steps — split the requested cache by architecture into UB, L1 and L0C; insert workspace (because UB and L1/L0C cannot be copied between directly, so data used in both places has to exist twice); separate CV cores; synchronize across cores; have the two vector cores each handle half the data; and finally deal with the hardware special case that L0C cannot be written to directly. One of their measurements: hand-written FlashAttention at 44.9 microseconds, 22.28 microseconds after manually pipelining it.
5. How third parties see it
The most systematic external tracking comes from SemiAnalysis's InferenceX. They've been recording DeepSeek V4 1.6T inference performance across hardware from Day 0 through Day 43, with a dedicated section analyzing inference on Ascend 950DT — how compute and communication overlap, and which compute streams Huawei used. One of their own observations is worth keeping: on the CUDA side, vLLM and SGLang were usable the day the model shipped, while ROCm wasn't usable on day one and only caught up on day 26, with an improvement of more than 100×. The ecosystem maturity gap is the genuinely hard part of this road. The same piece also notes that V4's fused MegaMoE kernel carries a theoretical 1.92× speedup over a naive implementation — working backwards, a naive implementation spends close to half its time in dispatch and combine.
Two industry figures as reference points: IDC's data has NVIDIA's share of China's AI accelerator market falling from 95% two years ago to 55% in 2025, with domestic chips passing 41% for the first time — Ascend alone shipping 812,000 cards. China Mobile's April AI supernode tender this year was for 6,208 cards at roughly 2.06 billion RMB, requiring CANN-ecosystem solutions across the board. Both are third-party figures; I'd re-check them before using them.
6. How far I got on my own machine
(After cloning the six repos locally, where the chain breaks. Green is what worked; red is what this machine doesn't have.)
The environment first: Python 3.11.15; /usr/local/Ascend does not exist; import torch_npu raises ModuleNotFoundError; which bisheng ccec finds nothing; there is no nvidia-smi command; and lspci shows zero NVIDIA, Ascend or Huawei devices. This is a pure CPU machine.
Then I actually tried to build DeepGEMM-Ascend. The system python3 doesn't even have setuptools, so it stopped at line 5. I made a clean venv, installed the dependencies, and the second attempt got to line 7:
import torch
ModuleNotFoundError: No module named 'torch'
The Ascend part only starts at line 28: ASCEND_HOME defaults to /usr/local/Ascend/ascend-toolkit/latest, then it imports torch_npu and links against ascendcl, hccl_fwk, torch_npu, opapi and dl. DeepEP-Ascend's CMakeLists is even more direct — under ${ASCEND_HOME_PATH}/aarch64-linux/ alone there are dozens of header directories, plus the clang/15.0.5 headers from the bisheng compiler.
The conclusion is that I never even reached the doorstep. Reproducing any of this yourself starts with owning an Ascend machine with CANN installed.
7. What's in the code but not in the announcement
(Distribution of Ascend-specific kernels in TileKernels. The quant directory is the largest: 12 files, 3,491 lines.)
TileKernels writes two implementations for each kernel. Take SwiGLU: the directory contains both swiglu_forward_kernel.py and swiglu_forward_asc.py. The Ascend one starts with import tilelang.ascend.language as T and from tilelang.ascend.language import simd as S, and its pipeline synchronization is hand-written:
T.ascend_set_flag('MTE3_MTE2', 0)
T.ascend_wait_flag('MTE2_V', input_slot)
The README's "backend selected automatically at runtime" is true, and the mechanism is an is_ascend() check in tile_kernels/config.py; the package can carry both backends in one environment. But a backend is not automatic translation — the Ascend implementation is written by hand. I counted the Ascend-specific files: quant 12 files / 3,491 lines, mhc 9 / 2,375, engram 6 / 969, moe 4 / 652, plus one each in transform and rand — 33 files, 7,729 lines in total.
Constraints scattered through asserts. A few: per-token quantization only supports groups of 32 channels; the E4M3 scale factor is CUDA-only; SwiGLU only supports sf_block=(1, 32); hyper-connections only support 4 streams; randn doesn't support graph capture on NPU because replay would reuse the same random seed; and norm's hidden dimension must be divisible by 128. The announcement's line that "every TileLang kernel has a corresponding high-performance implementation" holds — the price is that list of shapes and preconditions.
FlashMLA's Ascend kernel covers only 6 combinations. Under csrc/ascend_kernels/prefill/sparse there are 4 headers plus 6 script-generated instantiation files: V4.1, h64, prefill and decode, with and without sink, fp8 and fp4. Change the shape and you regenerate. This is not a general-purpose library.
There is no kernel file in DeepEP-Ascend with "ascend" in its name. Its kernels are organized by function under csrc/kernels/ep, comm, bucket, pp and engram — all .hpp, 47 files, 6,335 lines in total.
The same batch of commits carries a breaking change. FlashMLA removed Hopper support and V3, V3.2 and V4.0 in this release, changed the KV cache format, and requires switching back to the 15 September commit if you need older models or formats. The Ascend kernels here only serve V4.1. This version of the attention library moves with V4.1.
8. What this post doesn't do
I reproduced none of the performance numbers — I have no Ascend card. I did no independent verification of second-hand items like "160,000 chips" or the supernode timelines; each is marked with its source in the text. When the commercial HDK lands in mid-October, there will be a version a third party can actually check, and I'll come back and fill this in then.
References
- DeepSeek's announcement, "DeepSeek open-sources Ascend foundation components" (in Chinese): https://mp.weixin.qq.com/s/X41mKH4Ds-VXUAnK6M8Eww
- FlashMLA Ascend technical report (Chinese): https://github.com/deepseek-ai/FlashMLA/blob/20260930-ascend-open-source/docs/20260930-ascend-prefill-deep-dive-zh.md
- DeepGEMM-Ascend, DeepEP-Ascend, TileKernels, FlashMLA, DeepSelect: the same-named repositories under github.com/deepseek-ai
- TileLang main repo PR #3308: https://github.com/tile-ai/tilelang/pull/3308
- TileLang Ascend backend docs: https://github.com/tile-ai/tilelang/blob/main/tilelang/ascend/README.md
- Huawei, "Ascend 960 supernode launch" (2026-09-17): https://www.huawei.com/cn/news/2026/9/hc-ascend960-supernode
- Huawei supernode interconnect keynote transcript (2025-09, Atlas 950 specs): https://www.huawei.com/cn/news/2025/9/hc-xu-keynote-speech
- Ascend community, "TileLang AscendNPU IR from beginner to expert" (2026-05-08): https://www.hiascend.com/developer/techArticles/20260506-1
- Cailian Press, "Huawei's Atlas 950 hardware to appear at WAIC" (2026-07-07): https://www.cls.cn/detail/2418935
- Guancha, "DeepSeek reportedly to buy 160,000 Huawei Ascend 950DT chips" (2026-09-05): https://www.guancha.cn/economy/2026_09_05_830131.shtml
- Economic Observer, "Huawei computing power opens its chain" (2026-05-01): http://www.eeo.com.cn/2026/0501/860610.shtml
- SemiAnalysis InferenceX, "DeepSeekV4 1.6T Day 0 to Day 43": https://inferencex.semianalysis.com/blog/deepseekv4-16t-day-0-to-day-43-performance
Every number, error message and source in this post is in the list above; nothing here is repeated from somewhere else. I write one of these deep dives a week, published first in my own community in Chinese — this is the cleaned-up public version.
I also keep a small paid community where I write these engineering logs up in full, one deep dive a week. Its content is in Chinese — I mention it in case you read Chinese: AI落地实录 (¥25/year with the current new-member coupon, ¥50 list price).
Disclosure: this English version was translated and edited with an AI assistant from my own write-up. Every figure, error message and file count comes from my own runs or the linked primary sources — which is exactly why the parts I could not verify myself are marked as vendor claims.
Originally written in Chinese for my own notes; this English version is a translation.
Top comments (0)