We ran Qwen3-8B on Ascend (DashScope) and NVIDIA A40 (CUDA) with identical parameters across 61 prompts. The result: only 15% top-1 token agreement. 98% of prompts diverged.
Results:
- Top-1 token agreement: 15.02%
- Prompts diverging: 60/61 (98.4%)
- Top-5 overlap: 20.92%
For comparison, A40 vs RTX 6000 Ada (same CUDA stack) diverged on ~19% of prompts. A40 vs Ascend: 98%.
Full analysis: https://ruitong.io/posts/ascend-vs-cuda-equivalence
Code: https://github.com/loopeywho/ruitong
Top comments (0)