DEV Community

Shixin Zhang
Shixin Zhang

Posted on

Is Python Slower Than Julia? A TensorCircuit-NG Benchmark on H200

In scientific computing, Python is often associated with convenience but slow execution, while Julia is widely regarded as a language designed for high-performance computing. Before looking at the actual code or software stack, it can seem as though the performance ranking has already been decided.

But consider the following experiment. On the same NVIDIA H200, running the same variational quantum algorithm, TensorCircuit-NG takes about 0.86 seconds per evaluation, compared with 3.94 seconds for Yao.jl from the Julia ecosystem. The Python-facing TensorCircuit-NG implementation is therefore 4.6× faster in this particular benchmark.

This raises a broader question: how much can we really infer about scientific-computing performance from the programming language being used? In fields dominated by linear algebra, different languages share much of the same numerical infrastructure. And as AI becomes increasingly capable of translating, rewriting, and optimizing code across languages, some of our old intuitions about language performance are becoming misleading—and increasingly turning into biases.

First, what are the two implementations actually computing?

The benchmark comes from the variational quantum eigensolver (VQE). VQE represents a quantum state with a parameterized quantum circuit and iteratively adjusts the circuit parameters to minimize the energy. Each optimization step therefore requires repeated evaluations of the current energy and its gradient with respect to the circuit parameters.

The problem consists of 26 qubits, a 16-layer circuit, and 816 parameters, all in double precision. The physical model is the one-dimensional transverse-field Ising model. Both implementations start from the same uniform superposition state. Each layer first applies RZZ gates between neighboring qubits, followed by RX gates on every qubit.

Yao.jl uses a native GPU state-vector representation together with reverse-mode automatic differentiation. TensorCircuit-NG uses its JAX backend, with omeco to optimize the tensor-contraction order.

The resulting parameters agree element by element, and the maximum difference between the complete gradients is at the level of machine precision. The benchmark scripts, commands, and dependency versions are available in the Gist linked at the end of the article.

How much faster can it get—and how long do you have to wait?

Even for the same JAX computation, the way the computation is expressed can make a substantial difference.

The scan implementation represents the repeated circuit layers using a loop structure. The unrolled implementation explicitly expands all layers and lets the compiler optimize the resulting contraction graph as a whole. The two implementations illustrate a familiar trade-off between compilation time and execution time.

Implementation Initial evaluation-related time Energy + gradient after warm-up
TensorCircuit-NG/JAX, scan 11.7 s 0.864 s
TensorCircuit-NG/JAX, unrolled 206.9 s 0.396 s
Yao.jl/CUDA, state vector 29.0 s 3.938 s

The scan version has the shorter compilation overhead and subsequently runs in 0.864 seconds per evaluation. Fully unrolling the circuit roughly doubles the execution efficiency, reducing the time to about 0.396 seconds. Relative to the Yao.jl implementation, that is close to a 10× speedup.

The price is a much longer initial compilation: more than three minutes.

Using these numbers, the unrolled implementation becomes more efficient overall after roughly 400 evaluations. In a VQE optimization that may require thousands of iterations, paying 206 seconds for XLA JIT compilation in exchange for 0.396-second evaluations can therefore be a very favorable trade.

The key idea is straightforward: for a workload with enough repeated evaluations, fully exposing the circuit to the compiler allows XLA to aggressively fuse operations and optimize the entire computation graph.

What about Yao.jl's tensor-network implementation?

An obvious question is: Yao.jl now also has tensor-network tooling, so why is that not included in the main table? Is the comparison fair if the two implementations use different computational models?

We also tested YaoToEinsum, OMEinsum, and cuTENSOR in various combinations on the same H200.

The GPU tensor contraction itself works, and differentiation with respect to the input tensors can also be performed successfully. However, when YaoToEinsum converts the circuit into a tensor network, it uses in-place mutation. This causes reverse-mode differentiation to fail at that point, so the energy gradient with respect to the circuit parameters cannot be computed through the existing interface.

In other words, for the Julia software stack tested here, this route cannot currently complete the full task required by the benchmark. We therefore use the state-vector implementation in the main table because it is the Yao.jl configuration that can successfully perform the complete energy-and-gradient computation.

What if we only calculate the energy and do not require gradients?

We also tested the forward pass at the same problem size on the H200. Both implementations first use tensor networks to obtain the full state vector and then calculate the energy using their respective native interfaces. The entire process is included in the timing.

TensorCircuit-NG uses the same unrolled configuration as above, while Yao.jl uses its optimized contraction path. After warm-up, we repeated the calculation 10 times and compared the median execution time.

The median time per evaluation was 0.095 seconds for TensorCircuit-NG and 0.192 seconds for Yao.jl, with identical energy results.

So even for this forward-only calculation, where the tensor-network route in Yao.jl works successfully, TensorCircuit-NG takes roughly half the time.

If Python is not doing the heavy lifting, who is?

The most expensive tensor computations in TensorCircuit-NG are handed off to JAX and XLA.

Python specifies the model, the circuit, and the differentiation target. JAX transforms the computation into a graph, and XLA compiles that graph for execution on the GPU. The actual heavy numerical work does not run through the Python interpreter.

A large fraction of scientific computing ultimately consists of matrix multiplications, factorizations, tensor contractions, and related numerical kernels. On CPUs, different programming languages can rely on BLAS and LAPACK. On NVIDIA GPUs, they can use CUDA and its linear-algebra and tensor-computing libraries. Compilers can also generate or fuse specialized kernels.

None of this performance belongs exclusively to one programming language.

Julia can of course compile code and use GPUs as well. At that point, what determines performance is often much further down the stack: contraction order, data layout, memory traffic, kernel fusion, and the design of the higher-level interfaces and algorithms.

Even when two programs ultimately call the same low-level library, moving data an extra time or materializing a few unnecessarily large tensors can make a substantial difference in both runtime and memory consumption.

The difference between the scan and unrolled versions of TensorCircuit-NG already makes the point. The language did not change at all. What changed was the way the computation was organized. Compilation time and execution speed changed dramatically as a result.

The performance difference is therefore not fundamentally a difference between Python and Julia.

In the AI era, is it still worth taking sides over programming languages?

In the past, switching languages often meant switching an entire ecosystem of tools, workflows, and debugging habits. Familiarity with a language was itself a major productivity advantage, and the cost of moving to another language could be substantial.

That cost is shrinking rapidly.

AI can read existing code, modify configurations, rewrite kernels, and move between Python, Julia, C++, Rust, and CUDA. Combined with automated testing and performance measurements, it can iteratively optimize an implementation across language boundaries.

We can increasingly give an AI system a complete computational task rather than insisting that every part of the solution remain in the language we happen to know best—or like the most.

A programming language is an API for expressing computations at a particular level of abstraction. It is not a silver bullet for solving the underlying problem.

The numerical libraries are shared. The hardware is shared. And now increasingly capable AI tools can work across languages as well.

The programming-language fan economy is therefore starting to lose its relevance.

The TensorCircuit-NG benchmark presented here is simply one example of Python outperforming Julia. It does not establish that Python is faster than Julia in general. What it does illustrate is a much simpler first principle:

Put your effort into making the program good, rather than into proving that your favorite language is superior.

Top comments (0)