<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Shixin Zhang</title>
    <description>The latest articles on DEV Community by Shixin Zhang (@refractionray).</description>
    <link>https://dev.to/refractionray</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3763205%2F5c5af020-0b22-4443-aa68-28b3150f48e4.png</url>
      <title>DEV Community: Shixin Zhang</title>
      <link>https://dev.to/refractionray</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/refractionray"/>
    <language>en</language>
    <item>
      <title>From “Shifting Gears” to “Borrowing”: Entanglement Spectra Are Becoming a Language for Nonequilibrium Dynamics</title>
      <dc:creator>Shixin Zhang</dc:creator>
      <pubDate>Wed, 23 Sep 2026 03:30:29 +0000</pubDate>
      <link>https://dev.to/refractionray/from-shifting-gears-to-borrowing-entanglement-spectra-are-becoming-a-language-for-26p9</link>
      <guid>https://dev.to/refractionray/from-shifting-gears-to-borrowing-entanglement-spectra-are-becoming-a-language-for-26p9</guid>
      <description>&lt;p&gt;On September 6, we uploaded &lt;a href="https://arxiv.org/abs/2609.06643" rel="noopener noreferrer"&gt;&lt;em&gt;Entanglement Growth as Transport Across Schmidt Scales&lt;/em&gt;&lt;/a&gt; to arXiv. The central idea was a simple shift in perspective: when studying entanglement dynamics, it is not enough to ask &lt;strong&gt;how much entanglement has grown&lt;/strong&gt;. We should also ask &lt;strong&gt;which Schmidt scales carry the dominant spectral weight, and how that weight moves over time&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;While working on that paper, two natural questions kept coming back to us.&lt;/p&gt;

&lt;p&gt;Can transport across Schmidt scales be connected to a more direct and operational quantum-information task? And how does the slow diffusion of a conserved charge reshape this process?&lt;/p&gt;

&lt;p&gt;Those questions became the starting point of our new work.&lt;/p&gt;

&lt;p&gt;Over the following ten days, three independent papers appeared on arXiv. On September 15, arXiv:2609.16935 studied nonlocal magic and spectral structure across the many-body localization crossover. On September 16, arXiv:2609.18691 showed that chaotic dynamics can generate transient entanglement embezzlement at intermediate times. On September 17, arXiv:2609.20951 connected nonlocal magic and entanglement embezzlement to many-body dynamics, including free-fermion dynamics. All three papers cited our earlier work.&lt;/p&gt;

&lt;p&gt;In little more than a week, &lt;strong&gt;Schmidt-spectrum structure, nonlocal magic, entanglement embezzlement, and many-body transport&lt;/strong&gt; began to converge around a common nonequilibrium perspective.&lt;/p&gt;

&lt;p&gt;A lively research direction is becoming visible.&lt;/p&gt;

&lt;p&gt;This unusually rapid sequence of developments also says something about the changing pace of research. New tools, including AI, are shortening the cycles of literature search, numerical implementation, and cross-checking. Ideas emerging in different places can meet each other much faster than before.&lt;/p&gt;

&lt;p&gt;The scientific questions still come from accumulated understanding, discussion, and judgment. What has changed is the feedback loop: the distance from a question to a concrete result can increasingly be measured in &lt;strong&gt;days rather than months&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That is the sense of urgency we felt while working on our new paper, &lt;em&gt;&lt;a href="https://arxiv.org/abs/2609.26362" rel="noopener noreferrer"&gt;Entanglement Embezzlement from Diffusive Hydrodynamics&lt;/a&gt;&lt;/em&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  From “Where Is the Entanglement?” to “What Can It Do?”
&lt;/h2&gt;

&lt;p&gt;Our previous work decomposed the Schmidt spectrum into exponentially growing rank windows and used these “doubling windows” to track whether the dominant spectral weight resides at low or high Schmidt rank.&lt;/p&gt;

&lt;p&gt;This revealed something that the entanglement entropy alone does not show: &lt;strong&gt;entropy growth, spectral reshaping, and the migration of dominant weight across Schmidt scales do not necessarily happen at the same time.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The new work asks a more operational question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Can this motion in the Schmidt spectrum actually be turned into a useful quantum-information task?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The answer is &lt;strong&gt;entanglement embezzlement&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Suppose Alice and Bob share a high-dimensional entangled state. Using only local operations and classical communication, they can “borrow” several Bell pairs from this state while leaving the original state globally almost unchanged. The high-dimensional entangled state therefore acts as a kind of catalyst.&lt;/p&gt;

&lt;p&gt;How much entanglement can be borrowed depends not only on the total amount of entanglement, but also on &lt;strong&gt;how broadly the Schmidt weight is distributed along the logarithmic Schmidt-rank axis&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Borrowing a maximally entangled state of dimension $d$ effectively shifts the target spectrum by $\log d$ along this axis.&lt;/p&gt;

&lt;p&gt;If the original spectrum is sufficiently broad, this shift represents only a small relative displacement, and the borrowed state can remain close to the original one.&lt;/p&gt;

&lt;p&gt;In this way, the qualitative notion of &lt;strong&gt;transport across Schmidt scales&lt;/strong&gt; from our previous paper becomes a directly quantifiable quantum-information task: one can ask how many Bell pairs can be borrowed at a specified error tolerance.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The Schmidt spectrum viewed as a distribution over logarithmic Schmidt rank, and the spectral width that controls entanglement embezzlement.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Can a Conserved Charge Act as an Entanglement Account?
&lt;/h2&gt;

&lt;p&gt;To understand the mechanism quantitatively, consider a random pure state in a fixed global $U(1)$ charge sector. Such states provide a useful model for the equilibrium states of chaotic particle-number-conserving systems.&lt;/p&gt;

&lt;p&gt;The total particle number is fixed, but the particle number inside a subsystem fluctuates. Different subsystem charge sectors therefore occupy different regions of the Schmidt spectrum.&lt;/p&gt;

&lt;p&gt;Away from half filling, an effective &lt;strong&gt;charge bias&lt;/strong&gt; converts ordinary particle-number fluctuations into a corresponding width on the logarithmic Schmidt-rank axis.&lt;/p&gt;

&lt;p&gt;This gives us a finite-error conversion law: at a fixed error tolerance, the amount of entanglement that can be borrowed grows with system size according to a definite scaling law.&lt;/p&gt;

&lt;p&gt;The result goes beyond the usual asymptotic question of whether embezzlement is possible. It gives a quantitative finite-size relation and, for arbitrary target dimension, an optimal fidelity curve.&lt;/p&gt;

&lt;p&gt;The most striking contrast appears at half filling.&lt;/p&gt;

&lt;p&gt;There, particle-hole symmetry makes the effective charge bias vanish. The system can still possess a maximal volume-law entanglement entropy, yet the amount of entanglement that can be borrowed goes to zero.&lt;/p&gt;

&lt;p&gt;This cleanly separates two notions that are often treated as nearly synonymous:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;having a lot of entanglement is not the same as having entanglement that is easy to borrow.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The distinction is invisible if one looks only at the entropy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Entropy Has Already Arrived. Diffusion Is Still Catching Up.
&lt;/h2&gt;

&lt;p&gt;The equilibrium analysis tells us where the system eventually ends up. The dynamical question is how it gets there.&lt;/p&gt;

&lt;p&gt;Consider a one-dimensional $U(1)$-conserving random quantum circuit. The leading volume-law entanglement entropy develops rapidly and saturates on a ballistic timescale $t \sim L$, while complete relaxation of the conserved charge requires the much longer diffusive timescale&lt;/p&gt;

&lt;p&gt;$$&lt;br&gt;
t_{\mathrm{diff}} \sim L^2.&lt;br&gt;
$$&lt;/p&gt;

&lt;p&gt;This creates a parametrically broad intermediate regime in which the entanglement entropy already looks saturated, while the internal structure of the Schmidt spectrum is still evolving.&lt;/p&gt;

&lt;p&gt;This is precisely where entanglement embezzlement becomes useful.&lt;/p&gt;

&lt;p&gt;Diffusion causes the variance of the subsystem charge to grow as&lt;/p&gt;

&lt;p&gt;$$&lt;br&gt;
\mathrm{Var}(Q) \sim t^{1/2},&lt;br&gt;
$$&lt;/p&gt;

&lt;p&gt;so the standard deviation grows as&lt;/p&gt;

&lt;p&gt;$$&lt;br&gt;
\sigma_Q \sim t^{1/4}.&lt;br&gt;
$$&lt;/p&gt;

&lt;p&gt;The charge bias then maps this fluctuation width onto a width in logarithmic Schmidt rank. Under the conditions that the charge envelope is sufficiently smooth and that each charge sector is internally well mixed, the theory predicts that the amount of borrowable entanglement in the intermediate-time regime grows as&lt;/p&gt;

&lt;p&gt;$$&lt;br&gt;
E_{\mathrm{embezzle}} \sim t^{1/4}.&lt;br&gt;
$$&lt;/p&gt;

&lt;p&gt;Large-scale two-replica tensor-network simulations give a numerical slope close to this theoretical prediction, together with a consistent finite-size scaling governed by the diffusive timescale.&lt;/p&gt;

&lt;p&gt;The $1/4$ exponent has a simple physical origin. It comes from a chain of transformations:&lt;/p&gt;

&lt;p&gt;$$&lt;br&gt;
\text{charge diffusion}&lt;br&gt;
\;\longrightarrow\;&lt;br&gt;
\text{charge variance}&lt;br&gt;
\;\longrightarrow\;&lt;br&gt;
\text{fluctuation width}&lt;br&gt;
\;\longrightarrow\;&lt;br&gt;
\text{Schmidt-spectrum width}&lt;br&gt;
\;\longrightarrow\;&lt;br&gt;
\text{borrowable entanglement}.&lt;br&gt;
$$&lt;/p&gt;

&lt;p&gt;More explicitly, diffusion produces a $t^{1/2}$ growth of the charge variance. Taking the square root gives a $t^{1/4}$ growth of the fluctuation width. That width determines how large a spectral translation can be accommodated at fixed error, and therefore how much entanglement can be borrowed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Entropy has already arrived. Diffusion is still catching up.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Entanglement embezzlement provides a way to see this hidden late-stage dynamics that the entropy has already forgotten.&lt;/p&gt;

&lt;h2&gt;
  
  
  A New Phase of Research Measured in Days
&lt;/h2&gt;

&lt;p&gt;The cluster of papers appearing over the past two weeks approaches the problem from different models and different physical questions, but they point toward a common object: &lt;strong&gt;the structure of the entanglement spectrum as a dynamical coordinate&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Schmidt scales are beginning to provide a bridge between quantum resources and nonequilibrium many-body dynamics.&lt;/p&gt;

&lt;p&gt;Our new work takes the next step along this direction. Instead of treating motion in the Schmidt spectrum merely as a diagnostic, we ask what physical resource that motion actually makes available.&lt;/p&gt;

&lt;p&gt;How much of the spectral structure can be turned into borrowable entanglement?&lt;/p&gt;

&lt;p&gt;This turns a family of spectral diagnostics into an operational question in quantum information, and suggests that entanglement embezzlement can serve as a probe of dynamical structure that is invisible to conventional entanglement measures.&lt;/p&gt;

&lt;p&gt;There is also something broader happening here.&lt;/p&gt;

&lt;p&gt;AI and other new tools do not create scientific questions by themselves. But once an idea is already in motion, they can make it move much faster: literature can be searched more quickly, calculations can be implemented sooner, numerical results can be cross-checked almost immediately, and related ideas emerging elsewhere can be connected with much less delay.&lt;/p&gt;

&lt;p&gt;The result is a different research tempo.&lt;/p&gt;

&lt;p&gt;Ideas that are already on the road can now accelerate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It really does feel faster—and there is a sense that it is urging you to keep running.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Shi-Xin Zhang, Shuo Liu, and Yu-Qin Chen, “Entanglement Growth as Transport Across Schmidt Scales,” &lt;a href="https://arxiv.org/abs/2609.06643" rel="noopener noreferrer"&gt;arXiv:2609.06643&lt;/a&gt; (2026).&lt;/li&gt;
&lt;li&gt;Shan-Zhong Li and Zhi Li, “Nonlocal Magic across the Many-Body Localization Crossover,” &lt;a href="https://arxiv.org/abs/2609.16935" rel="noopener noreferrer"&gt;arXiv:2609.16935&lt;/a&gt; (2026).&lt;/li&gt;
&lt;li&gt;Matias Karjula, Teemu Ojanen, Kim Pöyhönen, Tapio Ala-Nissila, and Moein N. Ivaki, “Universal Entanglement Embezzlement and Divergent Nonlocal Magic from Generic Local Chaotic Quantum Evolution,” &lt;a href="https://arxiv.org/abs/2609.18691" rel="noopener noreferrer"&gt;arXiv:2609.18691&lt;/a&gt; (2026).&lt;/li&gt;
&lt;li&gt;Sreemayee Aditya, Piotr Sierant, and Xhek Turkeshi, “Nonlocal Magic Spreading in Many-body Quantum Dynamics: From Chaotic Evolution to Quasi-particle Picture in Integrable Models,” &lt;a href="https://arxiv.org/abs/2609.20951" rel="noopener noreferrer"&gt;arXiv:2609.20951&lt;/a&gt; (2026).&lt;/li&gt;
&lt;li&gt;Shi-Xin Zhang, Shuo Liu, and Yu-Qin Chen, “Entanglement Embezzlement from Diffusive Hydrodynamics,” &lt;a href="https://arxiv.org/abs/2609.26362" rel="noopener noreferrer"&gt;arXiv:2609.26362&lt;/a&gt; (2026).&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>quantum</category>
    </item>
    <item>
      <title>Chasing Quantum States Through Time: A Tour of Time-Evolution Methods in TensorCircuit-NG</title>
      <dc:creator>Shixin Zhang</dc:creator>
      <pubDate>Mon, 14 Sep 2026 01:35:26 +0000</pubDate>
      <link>https://dev.to/refractionray/chasing-quantum-states-through-time-a-tour-of-time-evolution-methods-in-tensorcircuit-ng-37ea</link>
      <guid>https://dev.to/refractionray/chasing-quantum-states-through-time-a-tour-of-time-evolution-methods-in-tensorcircuit-ng-37ea</guid>
      <description>&lt;p&gt;Flip a single spin in a spin chain and, at first glance, almost nothing seems to happen. It is just a small disturbance dropped into a large quantum system.&lt;/p&gt;

&lt;p&gt;But then the disturbance propagates. Magnetization changes. Entanglement grows. Observables evolve. If a simulation is to follow the system from beginning to end, it has to keep track of the quantum state as it moves through time.&lt;/p&gt;

&lt;p&gt;The Schrödinger equation itself looks deceptively simple&lt;/p&gt;

&lt;p&gt;One line on paper turns into a surprisingly diverse collection of numerical methods in code. Some methods keep the Hamiltonian and the full state space explicitly. Some break time evolution into small steps. Some never construct the Hamiltonian matrix at all, and only ask what the Hamiltonian does to a state. Others change the representation of the quantum state itself.&lt;/p&gt;

&lt;p&gt;For (n) qubits, a full state vector already contains (2^n) complex amplitudes. Before the state has evolved very far, memory may become the limiting factor.&lt;/p&gt;

&lt;p&gt;Different time-evolution algorithms therefore make different trade-offs between computational cost, memory, accuracy, and the complexity of the state representation.&lt;/p&gt;

&lt;p&gt;TensorCircuit-NG supports all of these approaches.&lt;/p&gt;




&lt;h2&gt;
  
  
  Solve Everything at Once: Exact Diagonalization
&lt;/h2&gt;

&lt;p&gt;Exact diagonalization takes the most straightforward approach: keep the Hamiltonian and the complete Hilbert space.&lt;/p&gt;

&lt;p&gt;For a time-independent Hamiltonian,&lt;/p&gt;

&lt;p&gt;$$&lt;br&gt;
|\psi(t)\rangle=e^{-iHt}|\psi(0)\rangle,&lt;br&gt;
$$&lt;/p&gt;

&lt;p&gt;and once (H) has been diagonalized, changing the time only requires applying phases to its eigencomponents. Evaluating the state at different times is therefore essentially independent of the Hilbert-space dimension after the initial diagonalization.&lt;/p&gt;

&lt;p&gt;The price is obvious: storing a dense Hamiltonian is extremely expensive.&lt;/p&gt;

&lt;p&gt;For example, a dense (16)-qubit Hamiltonian in complex128 already requires roughly 64 GB of memory. The actual diagonalization requires several times more working memory. The bill is paid upfront, but for small systems and simulations requiring many time points, that can be a very good trade.&lt;/p&gt;




&lt;h2&gt;
  
  
  Split Time into Pieces: Trotter, TEBD, and ODE Solvers
&lt;/h2&gt;

&lt;p&gt;Trotter decomposition takes a different route. Instead of applying the full evolution operator at once, it divides the total evolution time into small intervals and applies local evolution operators sequentially.&lt;/p&gt;

&lt;p&gt;The smaller the time step, the smaller the Trotter error—but the more steps have to be performed.&lt;/p&gt;

&lt;p&gt;Trotter evolution does not require constructing the full Hamiltonian matrix, but it still stores the complete state vector.&lt;/p&gt;

&lt;p&gt;TEBD goes one step further.&lt;/p&gt;

&lt;p&gt;Instead of representing the state as a dense vector, TEBD represents it as a matrix product state (MPS). Local gates are applied sequentially, and the bond dimension is truncated after each update.&lt;/p&gt;

&lt;p&gt;For a one-dimensional system, the memory requirement can scale only linearly with system size for fixed bond dimension, rather than exponentially with the number of sites. This makes simulations involving hundreds or even thousands of qubits possible—provided that entanglement does not drive the required bond dimension too high.&lt;/p&gt;

&lt;p&gt;The trade-off is now twofold: Trotter error from discretizing time, and truncation error from compressing the quantum state.&lt;/p&gt;

&lt;p&gt;ODE-based evolution takes yet another approach. Instead of explicitly constructing a product formula, it integrates the Schrödinger equation directly. This is particularly convenient for time-dependent Hamiltonians and driven systems.&lt;/p&gt;

&lt;p&gt;The number of integration steps is determined by the numerical integrator, often adaptively, with the target error controlled by the solver.&lt;/p&gt;




&lt;h2&gt;
  
  
  No Matrix Required: Krylov, Chebyshev, and &lt;code&gt;expm-multiply&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;For a time-independent Hamiltonian, the formal solution is&lt;/p&gt;

&lt;p&gt;$$&lt;br&gt;
|\psi(t)\rangle=e^{-iHt}|\psi(0)\rangle.&lt;br&gt;
$$&lt;/p&gt;

&lt;p&gt;But there is no reason to explicitly construct (e^{-iHt}).&lt;/p&gt;

&lt;p&gt;Krylov methods approximate the evolution in a small subspace generated by repeated applications of (H) to the initial state.&lt;/p&gt;

&lt;p&gt;Chebyshev methods approximate the exponential using a polynomial expansion.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;expm-multiply&lt;/code&gt; uses a scaled Taylor-series approach to approximate the action of the matrix exponential on a vector.&lt;/p&gt;

&lt;p&gt;The common idea is simple:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;We do not need the matrix exponential. We only need its action on the state.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;These methods can work directly with a matrix-vector-product (MVP) interface. The implementation only needs to answer the question&lt;/p&gt;

&lt;p&gt;$$&lt;br&gt;
|\phi\rangle \mapsto H|\phi\rangle.&lt;br&gt;
$$&lt;/p&gt;

&lt;p&gt;The Hamiltonian itself does not necessarily need to exist as an explicit matrix.&lt;/p&gt;

&lt;p&gt;This is particularly useful for large sparse or structured many-body Hamiltonians. Compared with explicitly constructing and multiplying sparse matrices, a specialized MVP implementation can avoid substantial memory traffic and exploit the structure of the underlying operators.&lt;/p&gt;

&lt;h3&gt;
  
  
  TenCirPauli: Extending the MVP Interface
&lt;/h3&gt;

&lt;p&gt;TensorCircuit-NG already provides MVP implementations for spin operators and Pauli Hamiltonians.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/tensorcircuit/TenCirPauli" rel="noopener noreferrer"&gt;TenCirPauli&lt;/a&gt; extends this abstraction to fermionic, bosonic, spin, and mixed many-body Hamiltonians, while also supporting MVPs after symmetry reduction.&lt;/p&gt;

&lt;p&gt;For example, particle-number conservation or total-magnetization conservation can be used to restrict the calculation to a symmetry sector. The resulting reduced operator can then be passed directly to an ODE solver, Krylov evolution, Chebyshev expansion, or &lt;code&gt;expm-multiply&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The evolution algorithm does not need to know where the operator came from.&lt;/p&gt;

&lt;p&gt;It only needs an MVP.&lt;/p&gt;

&lt;p&gt;This separation between &lt;strong&gt;operator construction&lt;/strong&gt; and &lt;strong&gt;operator action&lt;/strong&gt; is useful both numerically and architecturally: the same evolution machinery can operate on very different physical representations.&lt;/p&gt;




&lt;h2&gt;
  
  
  TDVP: Let Parameters Move Instead of the Full State
&lt;/h2&gt;

&lt;p&gt;Eventually, the full wavefunction becomes too large to store.&lt;/p&gt;

&lt;p&gt;At that point, another possibility is to stop representing the quantum state explicitly.&lt;/p&gt;

&lt;p&gt;Instead, represent it using a parameterized ansatz,&lt;/p&gt;

&lt;p&gt;$$&lt;br&gt;
|\psi(\boldsymbol{\theta})\rangle,&lt;br&gt;
$$&lt;/p&gt;

&lt;p&gt;and evolve the parameters (\boldsymbol{\theta}).&lt;/p&gt;

&lt;p&gt;Time-dependent variational principle (TDVP) projects the exact Schrödinger dynamics onto the tangent space of the chosen variational manifold. The simulation therefore follows the best direction available within the chosen ansatz rather than the full Hilbert space.&lt;/p&gt;

&lt;p&gt;Fewer parameters can mean dramatically lower computational cost—but also potentially larger projection error.&lt;/p&gt;

&lt;p&gt;TensorCircuit-NG supports several variants of this idea:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Variational-circuit TDVP&lt;/strong&gt; uses the output of a parameterized quantum circuit as the state representation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MPS-TDVP&lt;/strong&gt; uses an MPS as the variational manifold. For low-entanglement one-dimensional systems, it can reach system sizes far beyond dense-state simulation and is often more natural than TEBD for long-range interactions.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The key idea is the same:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Instead of following every amplitude in Hilbert space, follow the coordinates of a much smaller manifold.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  A Quick Comparison
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Method&lt;/th&gt;
&lt;th&gt;Time cost&lt;/th&gt;
&lt;th&gt;Memory&lt;/th&gt;
&lt;th&gt;Typical scale&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Exact diagonalization&lt;/td&gt;
&lt;td&gt;Essentially constant per additional time point after diagonalization&lt;/td&gt;
&lt;td&gt;Full Hamiltonian + full state&lt;/td&gt;
&lt;td&gt;~14 qubits without symmetry reduction; 16+ with symmetry&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Trotter / ODE&lt;/td&gt;
&lt;td&gt;Roughly linear in the number of time steps&lt;/td&gt;
&lt;td&gt;Full state&lt;/td&gt;
&lt;td&gt;~25–30 qubits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Krylov / Chebyshev / &lt;code&gt;expm-multiply&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Roughly linear in the number of evolution steps / MVP evaluations&lt;/td&gt;
&lt;td&gt;Full state + work vectors&lt;/td&gt;
&lt;td&gt;~25–30 qubits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Variational-circuit TDVP&lt;/td&gt;
&lt;td&gt;Roughly linear in evolution steps&lt;/td&gt;
&lt;td&gt;State + variational parameters&lt;/td&gt;
&lt;td&gt;~20 qubits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TEBD / MPS-TDVP&lt;/td&gt;
&lt;td&gt;Roughly linear in system size for fixed bond dimension&lt;/td&gt;
&lt;td&gt;MPS&lt;/td&gt;
&lt;td&gt;Hundreds to thousands of qubits in 1D&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These numbers are not hard limits. The actual boundary depends heavily on hardware, precision, Hamiltonian structure, entanglement growth, bond dimension, and how much memory one is willing to spend.&lt;/p&gt;

&lt;p&gt;And these days, that last variable can get expensive very quickly.&lt;/p&gt;




&lt;h2&gt;
  
  
  There Are More Ways to Evolve a Quantum State
&lt;/h2&gt;

&lt;p&gt;Some physical problems admit even more specialized representations.&lt;/p&gt;

&lt;h3&gt;
  
  
  Free-fermion dynamics
&lt;/h3&gt;

&lt;p&gt;Free-fermion evolution is one such special case.&lt;/p&gt;

&lt;p&gt;When both the Hamiltonian and initial state preserve fermionic Gaussian structure, TensorCircuit-NG's &lt;code&gt;FGSSimulator&lt;/code&gt; can evolve the system directly on the Gaussian-state manifold rather than storing the full many-body wavefunction.&lt;/p&gt;

&lt;h3&gt;
  
  
  Neural quantum states
&lt;/h3&gt;

&lt;p&gt;For more complicated entanglement structures, neural quantum states (NQS) combined with time-dependent variational Monte Carlo provide another route. Conceptually, NQS-tVMC is again a TDVP calculation, but with a neural-network wavefunction and Monte Carlo sampling replacing an explicitly stored state.&lt;/p&gt;

&lt;h3&gt;
  
  
  PEPS
&lt;/h3&gt;

&lt;p&gt;For genuinely two-dimensional systems, PEPS provides a natural extension of tensor-network representations. Real-time PEPS evolution remains substantially more difficult than its one-dimensional counterparts, but recent PEPS-tVMC approaches are beginning to combine PEPS representations with variational Monte Carlo.&lt;/p&gt;

&lt;p&gt;TensorCircuit-NG does not currently provide ready-to-run examples for all of these approaches. But the underlying infrastructure is designed to make such extensions possible.&lt;/p&gt;

&lt;p&gt;If your Agent needs another one, bring it along.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Algorithm Is Only Half the Story
&lt;/h2&gt;

&lt;p&gt;Time-evolution algorithms look quite similar on paper. In practice, the way they are executed can make a substantial difference.&lt;/p&gt;

&lt;p&gt;TensorCircuit-NG brings automatic differentiation, JIT compilation, vectorization, and GPU execution into the same computational workflow. This is particularly useful when quantum dynamics is not an isolated simulation, but part of a larger machine-learning or optimization pipeline.&lt;/p&gt;

&lt;h3&gt;
  
  
  JIT compilation
&lt;/h3&gt;

&lt;p&gt;Repeated evolution steps can be compiled into an optimized computation graph instead of being dispatched one operation at a time from Python.&lt;/p&gt;

&lt;h3&gt;
  
  
  GPU and &lt;code&gt;vmap&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;State-vector operations, MVPs, tensor contractions, and other numerical kernels can run directly on GPUs.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;vmap&lt;/code&gt; makes it possible to evolve batches of initial states or evaluate multiple time points in parallel when the problem structure allows it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Automatic differentiation
&lt;/h3&gt;

&lt;p&gt;Parameters do not have to be fixed inputs.&lt;/p&gt;

&lt;p&gt;Hamiltonian couplings, external fields, evolution times, and even initial-state parameters can all become differentiable quantities.&lt;/p&gt;

&lt;p&gt;This turns problems such as quantum-control optimization, Hamiltonian inverse design, pulse optimization, and target-state preparation into gradient-based computational problems rather than brute-force parameter sweeps.&lt;/p&gt;

&lt;p&gt;That is where the distinction between a physics simulator and a computational framework becomes important.&lt;/p&gt;




&lt;h2&gt;
  
  
  So Which Method Should You Use?
&lt;/h2&gt;

&lt;p&gt;Ideally, you should not have to memorize this table.&lt;/p&gt;

&lt;p&gt;Given a concrete problem, the right choice depends on the Hamiltonian, system size, entanglement growth, time dependence, available hardware, desired accuracy, and whether gradients are required.&lt;/p&gt;

&lt;p&gt;This is also increasingly a job for an Agent.&lt;/p&gt;

&lt;p&gt;Give the Agent the problem, the TensorCircuit-NG documentation, and the relevant hardware constraints. It can decide whether the calculation calls for exact diagonalization, an MVP-based method, Trotter evolution, MPS, TDVP, or something more specialized—and tune the numerical parameters accordingly.&lt;/p&gt;

&lt;p&gt;There is a practical reason to explicitly tell an Agent about TensorCircuit-NG.&lt;/p&gt;

&lt;p&gt;If you do not, it may reach for one of the established "standard" packages by default. They are standards for good reasons.&lt;/p&gt;

&lt;p&gt;They are also very good at being standard.&lt;/p&gt;

&lt;p&gt;And sometimes, very standard means very slow.&lt;/p&gt;




&lt;h2&gt;
  
  
  Examples and References
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Exact evolution, Krylov, and ODE comparison&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://github.com/tensorcircuit/tensorcircuit-ng/blob/master/examples/time_evolution_comparison.py" rel="noopener noreferrer"&gt;https://github.com/tensorcircuit/tensorcircuit-ng/blob/master/examples/time_evolution_comparison.py&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Trotter decomposition&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://github.com/tensorcircuit/tensorcircuit-ng/blob/master/examples/timeevolution_trotter.py" rel="noopener noreferrer"&gt;https://github.com/tensorcircuit/tensorcircuit-ng/blob/master/examples/timeevolution_trotter.py&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;TEBD and MPS&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://github.com/tensorcircuit/tensorcircuit-ng/blob/master/examples/xyzmodel_tebd.py" rel="noopener noreferrer"&gt;https://github.com/tensorcircuit/tensorcircuit-ng/blob/master/examples/xyzmodel_tebd.py&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;ODE-based evolution with time-dependent driving&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://github.com/tensorcircuit/tensorcircuit-ng/blob/master/examples/ng_whitepaper/IVD_time_evolution.py" rel="noopener noreferrer"&gt;https://github.com/tensorcircuit/tensorcircuit-ng/blob/master/examples/ng_whitepaper/IVD_time_evolution.py&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Krylov time evolution&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://github.com/tensorcircuit/tensorcircuit-ng/blob/master/examples/krylov_time_evolution.py" rel="noopener noreferrer"&gt;https://github.com/tensorcircuit/tensorcircuit-ng/blob/master/examples/krylov_time_evolution.py&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Chebyshev evolution&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://github.com/tensorcircuit/tensorcircuit-ng/blob/master/examples/chebyshev_evol.py" rel="noopener noreferrer"&gt;https://github.com/tensorcircuit/tensorcircuit-ng/blob/master/examples/chebyshev_evol.py&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;&lt;code&gt;expm-multiply&lt;/code&gt; evolution&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://github.com/tensorcircuit/tensorcircuit-ng/blob/master/examples/expm_multiply_evol.py" rel="noopener noreferrer"&gt;https://github.com/tensorcircuit/tensorcircuit-ng/blob/master/examples/expm_multiply_evol.py&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Fermionic MVP time evolution&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://github.com/tensorcircuit/tensorcircuit-ng/blob/master/examples/tencirpauli_fermion_mvp_timeevolution.py" rel="noopener noreferrer"&gt;https://github.com/tensorcircuit/tensorcircuit-ng/blob/master/examples/tencirpauli_fermion_mvp_timeevolution.py&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;TenCirPauli: many-body operators, symmetry reduction, and MVP&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://github.com/tensorcircuit/TenCirPauli" rel="noopener noreferrer"&gt;https://github.com/tensorcircuit/TenCirPauli&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Variational-circuit TDVP&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://github.com/tensorcircuit/tensorcircuit-ng/blob/master/examples/variational_dynamics.py" rel="noopener noreferrer"&gt;https://github.com/tensorcircuit/tensorcircuit-ng/blob/master/examples/variational_dynamics.py&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;MPS-TDVP&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://github.com/tensorcircuit/tensorcircuit-ng/blob/master/examples/one_site_tdvp.py" rel="noopener noreferrer"&gt;https://github.com/tensorcircuit/tensorcircuit-ng/blob/master/examples/one_site_tdvp.py&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>quantum</category>
      <category>agents</category>
    </item>
    <item>
      <title>The Origin of Quantum Architecture Search</title>
      <dc:creator>Shixin Zhang</dc:creator>
      <pubDate>Fri, 11 Sep 2026 09:42:07 +0000</pubDate>
      <link>https://dev.to/refractionray/the-origin-of-quantum-architecture-search-1deo</link>
      <guid>https://dev.to/refractionray/the-origin-of-quantum-architecture-search-1deo</guid>
      <description>&lt;p&gt;Some technical terms become so familiar that we eventually forget that they had an origin.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quantum Architecture Search (QAS)&lt;/strong&gt; is now one of those terms.&lt;/p&gt;

&lt;p&gt;A strict exact-match search for the phrase &lt;em&gt;“Quantum Architecture Search”&lt;/em&gt; on Google Scholar now returns more than &lt;strong&gt;800 academic papers&lt;/strong&gt;. The phrase appears in papers on variational quantum algorithms, quantum machine learning, quantum compiling, reinforcement learning, hardware-aware circuit design, and many other topics.&lt;/p&gt;

&lt;p&gt;But the term itself has a rather precise starting point.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the name came from
&lt;/h2&gt;

&lt;p&gt;The phrase &lt;em&gt;Quantum Architecture Search&lt;/em&gt; was introduced in our work &lt;strong&gt;“Differentiable Quantum Architecture Search,”&lt;/strong&gt; first posted on arXiv on &lt;strong&gt;October 16, 2020&lt;/strong&gt; &lt;a href="https://arxiv.org/abs/2010.08561" rel="noopener noreferrer"&gt;https://arxiv.org/abs/2010.08561&lt;/a&gt;. The paper was subsequently published in &lt;em&gt;Quantum Science and Technology&lt;/em&gt; in 2022.&lt;/p&gt;

&lt;p&gt;The idea behind the name was straightforward.&lt;/p&gt;

&lt;p&gt;At the time, there were already discussions of searching for quantum ansätze and automatically designing variational circuits. The word &lt;em&gt;ansatz&lt;/em&gt;, however, naturally suggests a particular variational formulation: one chooses a parameterized family of quantum states or circuits and then optimizes its parameters.&lt;/p&gt;

&lt;p&gt;We wanted to describe something more general.&lt;/p&gt;

&lt;p&gt;The inspiration was partly the distinction that had already become familiar in machine learning between &lt;strong&gt;parameter optimization&lt;/strong&gt; and &lt;strong&gt;architecture search&lt;/strong&gt;. Neural Architecture Search does not merely optimize the weights of a neural network; it searches for the computational structure itself.&lt;/p&gt;

&lt;p&gt;We asked a similar question for quantum circuits:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Instead of optimizing a predefined quantum ansatz, can we search for the quantum architecture itself?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That was the reason for choosing the word &lt;strong&gt;architecture&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;It deliberately expanded the scope beyond &lt;em&gt;ansatz search&lt;/em&gt;. The object being searched could include the choice and arrangement of gates, circuit connectivity, depth, parameter sharing, and other structural degrees of freedom. In other words, the problem was not restricted to finding a better ansatz for one particular VQA.&lt;/p&gt;

&lt;p&gt;The term &lt;em&gt;Quantum Architecture Search&lt;/em&gt; was intended to name this broader problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  From a name to a research direction
&lt;/h2&gt;

&lt;p&gt;Our first work introduced &lt;strong&gt;Differentiable Quantum Architecture Search (DQAS)&lt;/strong&gt;, treating the architecture search problem in an end-to-end differentiable framework. The paper demonstrated the idea through several different circuit-design problems, including unitary decomposition, noise-aware circuit redesign, and automatically discovering circuit layouts for QAOA.&lt;/p&gt;

&lt;p&gt;The important point was not any particular search algorithm.&lt;/p&gt;

&lt;p&gt;It was the abstraction.&lt;/p&gt;

&lt;p&gt;Once quantum circuit design is viewed as an &lt;strong&gt;architecture search problem&lt;/strong&gt;, many different optimization techniques can naturally enter the picture: reinforcement learning, evolutionary algorithms, gradient-based methods, predictors, meta-learning, and others. The search strategy can change without changing the underlying problem definition.&lt;/p&gt;

&lt;p&gt;That abstraction turned out to be useful.&lt;/p&gt;

&lt;p&gt;Subsequent work adopted the terminology for a variety of settings which introduced differentiable, reinforcement-learning, predictor-based, distributed, hardware-aware, and other variants of QAS.&lt;/p&gt;

&lt;p&gt;A 2024 survey was eventually devoted specifically to the subject: &lt;strong&gt;“Quantum Architecture Search: A Survey.”&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;At that point, QAS was no longer simply the name of one method. It had become a recognizable research direction.&lt;/p&gt;

&lt;h2&gt;
  
  
  A useful historical distinction
&lt;/h2&gt;

&lt;p&gt;There is an interesting difference between inventing an algorithm and introducing a research vocabulary.&lt;/p&gt;

&lt;p&gt;An algorithm can be independently rediscovered. A useful vocabulary is different: once it captures a real conceptual boundary, it gives many subsequent works a common language.&lt;/p&gt;

&lt;p&gt;This is why the history of the phrase matters.&lt;/p&gt;

&lt;p&gt;There were certainly earlier studies on automatically optimizing quantum circuits, variational circuit structures, evolutionary circuit design, and related problems. But &lt;strong&gt;“Quantum Architecture Search” as a named research concept was introduced in the 2020 DQAS work&lt;/strong&gt;. The paper explicitly defined QAS as the automation of quantum-circuit architecture engineering.&lt;/p&gt;

&lt;p&gt;The choice of &lt;em&gt;architecture&lt;/em&gt; was therefore intentional. It was not simply another name for optimizing a variational ansatz.&lt;/p&gt;

&lt;p&gt;It was meant to define a broader abstraction:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;from searching for an ansatz to searching for a quantum architecture.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How far has the term traveled?
&lt;/h2&gt;

&lt;p&gt;The scale of adoption is perhaps the most interesting part of the story.&lt;/p&gt;

&lt;p&gt;A phrase that did not exist as an established research label in 2020 is now used across hundreds of papers and across multiple subfields of quantum computing. Recent work describes QAS as a “prominent paradigm” for quantum circuit design, and dedicated surveys and specialized QAS methods have appeared as the field has expanded.&lt;/p&gt;

&lt;p&gt;The original paper itself has also received substantial recognition. In 2025, &lt;em&gt;Differentiable Quantum Architecture Search&lt;/em&gt; was selected for the &lt;strong&gt;IOP Publishing China Top Cited Paper Award&lt;/strong&gt;. IOP's 2025 award covered research published by China-based corresponding authors across several disciplines; the award list included only &lt;strong&gt;eight papers in Physics&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That trajectory is a useful reminder that terminology is not merely cosmetic.&lt;/p&gt;

&lt;p&gt;Sometimes a new name identifies a new way of organizing a problem. If the abstraction is useful enough, other researchers begin building on it, the terminology becomes shared vocabulary, and eventually the vocabulary itself becomes part of the structure of the field.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Entanglement Doesn’t Just Grow. It Can Also Shift Gears: Four Clocks in Quantum Dynamics</title>
      <dc:creator>Shixin Zhang</dc:creator>
      <pubDate>Thu, 10 Sep 2026 05:16:55 +0000</pubDate>
      <link>https://dev.to/refractionray/entanglement-doesnt-just-grow-it-can-also-shift-gears-four-clocks-in-quantum-dynamics-1blk</link>
      <guid>https://dev.to/refractionray/entanglement-doesnt-just-grow-it-can-also-shift-gears-four-clocks-in-quantum-dynamics-1blk</guid>
      <description>&lt;p&gt;When we study quantum many-body dynamics, entanglement entropy is usually the first quantity we look at. It compresses an enormous quantum state into a single number, telling us when entanglement grows, how fast it grows, and where it eventually saturates. Over time, “increasing entanglement entropy” has almost become synonymous with “entanglement is developing.”&lt;/p&gt;

&lt;p&gt;But is &lt;em&gt;more&lt;/em&gt; really the whole story?&lt;/p&gt;

&lt;p&gt;Two states can have the same total amount of entanglement while having completely different internal structures. Think of two people holding the same amount of money: one keeps almost all of it in a single account, while the other spreads it across hundreds or thousands of accounts. The total is identical, but the distributions are not.&lt;/p&gt;

&lt;p&gt;The same is true for entanglement. Two quantum states with similar entanglement entropy can have very different Schmidt spectra. In one, a few Schmidt coefficients may still dominate. In the other, the weight may already have spread across a very large rank. Entanglement entropy sees the overall amount; the spectrum reveals &lt;em&gt;how that entanglement is organized&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;This changes the question we should ask.&lt;/p&gt;

&lt;p&gt;Instead of tracking only &lt;strong&gt;how much entanglement has been generated&lt;/strong&gt;, we can also ask &lt;strong&gt;what shape the entanglement spectrum is taking&lt;/strong&gt; and &lt;strong&gt;which Schmidt-rank scale carries the dominant weight&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Once we do this, what looked like a single process—“entanglement growth”—splits into several distinct dynamical events that do not necessarily happen at the same time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Opening Up a Number into a Spectrum
&lt;/h2&gt;

&lt;p&gt;Consider a bipartition of a pure quantum state. Its Schmidt decomposition expresses the state as a collection of correlated modes, each carrying a certain weight. Sorting these weights from largest to smallest gives the &lt;strong&gt;Schmidt spectrum&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Entanglement entropy compresses this entire spectrum into one number. The full spectrum contains much more information: which components dominate, how uneven the distribution is, and at what rank scale the major weight resides.&lt;/p&gt;

&lt;p&gt;To make the last feature visible, we introduce a simple notion of &lt;strong&gt;Schmidt-scale coordinate&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Imagine sliding a window through the ordered Schmidt spectrum. Starting at rank (r), we collect the total weight between (r) and (2r), and then ask which such window carries the largest weight. The weight of the winning window tells us how concentrated the spectrum is; its position tells us the &lt;strong&gt;dominant Schmidt scale&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;If the spectral head remains dominant, the system is still in a low-rank regime. If a window at higher rank eventually overtakes the head, the dominant scale has undergone a &lt;strong&gt;shift of gears&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Nothing is physically moving through real space here. What moves is the location, in the &lt;em&gt;rank-ordered spectrum&lt;/em&gt;, of where the dominant Schmidt weight resides.&lt;/p&gt;

&lt;h2&gt;
  
  
  Thermalization and MBL: Two Routes Through the Entanglement Spectrum
&lt;/h2&gt;

&lt;p&gt;A disordered spin chain provides a particularly clear contrast. Neighboring spins interact, while each site experiences a random local field.&lt;/p&gt;

&lt;p&gt;In the &lt;strong&gt;thermal regime&lt;/strong&gt;, the system gradually thermalizes. Entanglement entropy grows rapidly, with an approximately linear regime. At the same time, the largest Schmidt weights progressively lose their dominance, and the dominant spectral window moves toward higher ranks.&lt;/p&gt;

&lt;p&gt;Entanglement is not merely being generated. Its weight is being redistributed toward increasingly complex Schmidt scales.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;many-body localized (MBL)&lt;/strong&gt; regime tells a different story.&lt;/p&gt;

&lt;p&gt;Interactions still generate entanglement, and the entanglement entropy can continue to grow logarithmically. Higher-rank Schmidt components continue to appear as well. Yet the dominant spectral weight can remain pinned near the head of the spectrum, without undergoing the same shift toward higher rank scales.&lt;/p&gt;

&lt;p&gt;This is something entanglement entropy alone cannot reveal.&lt;/p&gt;

&lt;p&gt;Thermal and MBL systems can both exhibit entanglement growth, but the internal dynamics of their spectra can be fundamentally different. In particular, &lt;strong&gt;the generation of more entanglement and the transport of dominant weight across Schmidt scales can become decoupled&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure: Dynamics of different entanglement-related quantities in thermal and many-body-localized systems&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  One Spectrum, Four Clocks
&lt;/h2&gt;

&lt;p&gt;The Schmidt-scale coordinate is only one way to look inside the spectrum. The &lt;em&gt;shape&lt;/em&gt; of the spectrum is changing as well.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Anti-flatness&lt;/strong&gt; measures how uneven the Schmidt weights are. It vanishes when all nonzero Schmidt weights are equal, and becomes large when a few large weights coexist with many much smaller ones.&lt;/p&gt;

&lt;p&gt;Starting from a product state, there is initially only one nonzero Schmidt weight, so the anti-flatness is zero. As the system evolves, new Schmidt components appear, but they are highly unequal in magnitude. The spectrum therefore becomes increasingly uneven.&lt;/p&gt;

&lt;p&gt;Later, as weight spreads across more and more components, the distribution begins to flatten again. Anti-flatness consequently develops a pronounced intermediate-time peak—a temporary barrier in the dynamics.&lt;/p&gt;

&lt;p&gt;There is another quantity that follows its own clock: &lt;strong&gt;magic&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Magic is a resource-theoretic notion in quantum computation. Certain highly entangled states—stabilizer states, for example—can nevertheless be simulated efficiently on a classical computer. Nonlocal magic quantifies how far the evolving entanglement structure departs from this efficiently simulable stabilizer description.&lt;/p&gt;

&lt;p&gt;Its dynamics can develop an intermediate-time barrier of its own.&lt;/p&gt;

&lt;p&gt;We can therefore place four characteristic events on the same time axis:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The entanglement clock:&lt;/strong&gt; the peak in the entanglement-growth rate, around &lt;em&gt;[figure]&lt;/em&gt;;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The roughness clock:&lt;/strong&gt; the peak of anti-flatness, around &lt;em&gt;[figure]&lt;/em&gt;;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The magic clock:&lt;/strong&gt; the peak of the exact nonlocal magic, around &lt;em&gt;[figure]&lt;/em&gt;;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The shift clock:&lt;/strong&gt; the point at which more than half of the samples have left the spectral head, around &lt;em&gt;[figure]&lt;/em&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The important point is their ordering.&lt;/p&gt;

&lt;p&gt;Even after the entanglement-growth rate has peaked, the entanglement spectrum continues to reorganize. After the spectrum reaches its maximum roughness and nonlocal magic reaches its peak, the dominant spectral weight still takes additional time to migrate toward higher Schmidt scales.&lt;/p&gt;

&lt;p&gt;This ordering remains robust across system sizes and disorder strengths. In the MBL regime, however, the fourth clock may simply never ring.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure: The ordering of the four dynamical clocks&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does the Shift Clock Ring Last?
&lt;/h2&gt;

&lt;p&gt;The answer lies in the competition between the &lt;strong&gt;head&lt;/strong&gt; and the &lt;strong&gt;tail&lt;/strong&gt; of the Schmidt spectrum.&lt;/p&gt;

&lt;p&gt;Imagine that the largest Schmidt weight sits at the head, while the remaining weight forms a long tail. As weight flows from the head into the tail, different measures respond at different thresholds.&lt;/p&gt;

&lt;p&gt;When the tail carries roughly one quarter of the total weight, the contrast between the dominant head and the rest of the spectrum is strongest, producing the peak in spectral roughness.&lt;/p&gt;

&lt;p&gt;As the tail approaches roughly half of the total weight, the nonlocal magic reaches its maximum.&lt;/p&gt;

&lt;p&gt;Only when the tail grows sufficiently large—roughly beyond two thirds in the relevant spectral picture—can a higher-rank window overtake the original head and trigger the shift of the dominant Schmidt scale.&lt;/p&gt;

&lt;p&gt;In this sense, the ordering of &lt;strong&gt;roughness → magic → shift&lt;/strong&gt; is already encoded in the geometry of the spectrum itself.&lt;/p&gt;

&lt;p&gt;The same ordering also appears in random quantum circuits, suggesting that these clocks are not merely a peculiarity of one particular spin-chain model, but may reflect a more general feature of entanglement-spectrum dynamics.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Changes Is Not Just the Quantity We Measure, but the Question We Ask
&lt;/h2&gt;

&lt;p&gt;Each observable answers a different question:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Entanglement entropy:&lt;/strong&gt; How much entanglement is there?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Anti-flatness:&lt;/strong&gt; How uneven is the Schmidt spectrum?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Nonlocal magic:&lt;/strong&gt; How far has the entanglement structure moved away from a stabilizer description?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Schmidt-scale coordinate:&lt;/strong&gt; At what rank scale does the dominant weight reside?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The point of introducing four clocks is not simply to add four more curves to a plot of entanglement entropy.&lt;/p&gt;

&lt;p&gt;It is to separate processes that have traditionally been compressed into the single phrase &lt;strong&gt;“entanglement growth.”&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Entanglement is generated.&lt;/p&gt;

&lt;p&gt;The spectrum becomes rough.&lt;/p&gt;

&lt;p&gt;Nonlocal magic reaches a barrier.&lt;/p&gt;

&lt;p&gt;And eventually, the dominant weight shifts to a new Schmidt-rank scale.&lt;/p&gt;

&lt;p&gt;Entanglement entropy is an extraordinarily useful number. Perhaps it has become &lt;em&gt;too&lt;/em&gt; convenient: by reducing a complicated spectrum to a single scalar, it can make several distinct dynamical processes look like one.&lt;/p&gt;

&lt;p&gt;Once we open the spectrum back up, a hidden set of gears becomes visible.&lt;/p&gt;

&lt;p&gt;They do not all turn together.&lt;/p&gt;

&lt;p&gt;And the dynamics of entanglement, it turns out, is not simply a story of getting &lt;strong&gt;more&lt;/strong&gt; entangled. It is also a story of &lt;strong&gt;reorganization, redistribution, and changing scale&lt;/strong&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;References&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Lv Zhang, Shi-Xin Zhang, Heng Fan, and Shuo Liu, &lt;em&gt;Revealing Entanglement-Growth Mechanisms through the Magic Barrier&lt;/em&gt;, arXiv:2607.09875 (2026).&lt;/li&gt;
&lt;li&gt;Shi-Xin Zhang, Shuo Liu, and Yu-Qin Chen, &lt;em&gt;Entanglement Growth as Transport Across Schmidt Scales&lt;/em&gt;, arXiv:2609.06643 (2026).&lt;/li&gt;
&lt;/ol&gt;
&lt;/blockquote&gt;

</description>
      <category>quantum</category>
      <category>research</category>
    </item>
    <item>
      <title>Goodbye to Manual Code Review: The Vibe Coding Experiment Behind TenCirPauli</title>
      <dc:creator>Shixin Zhang</dc:creator>
      <pubDate>Wed, 26 Aug 2026 12:12:37 +0000</pubDate>
      <link>https://dev.to/refractionray/goodbye-to-manual-code-review-the-vibe-coding-experiment-behind-tencirpauli-478e</link>
      <guid>https://dev.to/refractionray/goodbye-to-manual-code-review-the-vibe-coding-experiment-behind-tencirpauli-478e</guid>
      <description>&lt;p&gt;In the previous article, we discussed why we started TenCirPauli and why we chose its technology stack. Many structured tasks in quantum operator processing involve dynamic data structures, discrete computations, and variable-length results. Providing a high-performance, Rust-native runtime for these workloads is the product motivation behind TenCirPauli.&lt;/p&gt;

&lt;p&gt;This article is about &lt;strong&gt;how it is developed&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;TenCirPauli is an extreme experiment in software development: &lt;strong&gt;the implementation is delegated entirely to AI. The developer does not write—or even read—any line of code.&lt;/strong&gt; The goal is to explore whether this new development paradigm can support a rigorous scientific computing software project at meaningful scale.&lt;/p&gt;

&lt;h2&gt;
  
  
  TensorCircuit-NG Still Uses the “Old-School” Approach
&lt;/h2&gt;

&lt;p&gt;TensorCircuit-NG takes a relatively conservative approach to development.&lt;/p&gt;

&lt;p&gt;When we encounter an issue or want to add a feature, we first work out the implementation strategy and then give the AI detailed instructions. After the AI makes the changes, a human reviews every modified file and essentially every significant addition or deletion, checking that the implementation follows the project's conventions and design principles.&lt;/p&gt;

&lt;p&gt;For a mature project with years of history and a large user base, this approach may be necessary.&lt;/p&gt;

&lt;p&gt;Such projects accumulate a large amount of &lt;strong&gt;implicit knowledge&lt;/strong&gt; that is never fully documented: why an interface is named a certain way, what seemingly unnecessary branch exists for backward compatibility, which coding patterns are intentional, and which internal conventions should never be changed casually. This knowledge is scattered across source code, issues, historical behavior, and the experience of maintainers. It is extremely difficult for an AI to reconstruct all of that context reliably.&lt;/p&gt;

&lt;p&gt;This development model has kept TensorCircuit-NG stable. It protects existing abstractions, semantics, and coding conventions, while preventing AI from introducing local workarounds simply to get a task done.&lt;/p&gt;

&lt;p&gt;If we simply started “vibe coding” inside such a project, however, the result could be very different. Old implicit conventions would gradually be overwritten, while new conventions would not have enough time to stabilize. Interfaces and internal structures could quickly lose coherence.&lt;/p&gt;

&lt;p&gt;That is why we chose a new project for this experiment.&lt;/p&gt;

&lt;p&gt;A new project has no historical baggage to preserve. Public interfaces, internal boundaries, testing methodology, and Agent collaboration workflows can all be designed from day one—and, importantly, they can be developed according to the AI's own preferences rather than being constrained by decades of accumulated human conventions.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Development Workflow
&lt;/h2&gt;

&lt;p&gt;TenCirPauli has already grown far beyond the scale of a toy demonstration. The project now contains &lt;strong&gt;more than 80,000 lines of code&lt;/strong&gt;, roughly half Python and half Rust.&lt;/p&gt;

&lt;p&gt;That is large enough for interface design, module boundaries, performance, and maintainability issues to emerge naturally. It is also large enough to test whether this development model can actually scale.&lt;/p&gt;

&lt;p&gt;In TenCirPauli, the developer's role is concentrated on &lt;strong&gt;goals, specifications, key architectural decisions, and acceptance criteria&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The workflow roughly looks like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A high-capability AI analyzes the goal and develops an initial Spec.&lt;/li&gt;
&lt;li&gt;The human decides on key interfaces, semantics, and architectural trade-offs, and aligns them with the Spec.&lt;/li&gt;
&lt;li&gt;A medium-capability AI implements the design.&lt;/li&gt;
&lt;li&gt;A high-capability AI reviews the implementation and generates a Review Report.&lt;/li&gt;
&lt;li&gt;The implementation AI fixes the identified issues.&lt;/li&gt;
&lt;li&gt;Automated tests and benchmarks provide objective evidence.&lt;/li&gt;
&lt;li&gt;The process is repeated as the project evolves.&lt;/li&gt;
&lt;li&gt;Periodic global AI inspections scan the repository for accumulated problems.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;In other words:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;High-capability AI&lt;/strong&gt; → formulate goals and write the Spec&lt;br&gt;
↓&lt;br&gt;
&lt;strong&gt;Human&lt;/strong&gt; → decide key interfaces, semantics, and architectural trade-offs&lt;br&gt;
↓&lt;br&gt;
&lt;strong&gt;Medium-capability AI&lt;/strong&gt; → implement the Spec&lt;br&gt;
↓&lt;br&gt;
&lt;strong&gt;High-capability AI&lt;/strong&gt; → review the implementation and generate a Review Report&lt;br&gt;
↓&lt;br&gt;
&lt;strong&gt;Medium-capability AI&lt;/strong&gt; → fix the issues&lt;/p&gt;

&lt;p&gt;The division of labor is deliberate.&lt;/p&gt;

&lt;p&gt;The high-capability model is responsible for understanding the problem, proposing solutions, and evaluating the result. The medium-capability model handles implementation and repairs. The human makes design trade-offs and keeps the overall system aligned. Automated infrastructure provides continuous evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Spec Is Where Human–AI Collaboration Matters Most
&lt;/h2&gt;

&lt;p&gt;A Spec is much more than a feature checklist.&lt;/p&gt;

&lt;p&gt;For a nontrivial scientific computing component, it needs to define the API surface and internal boundaries: invocation patterns, conceptual terminology, return values, data lifetimes, error types, the division of responsibilities between languages, performance and resource requirements, as well as independent reference implementations and benchmark-based acceptance criteria.&lt;/p&gt;

&lt;p&gt;Writing the initial Spec is therefore the most human-intensive part of the entire workflow.&lt;/p&gt;

&lt;p&gt;A high-capability AI can propose several interface designs, compare approaches used by existing libraries, and analyze complexity, resource consumption, language boundaries, and future extensibility.&lt;/p&gt;

&lt;p&gt;The developer, meanwhile, keeps asking the more fundamental questions:&lt;/p&gt;

&lt;p&gt;What does the user actually need?&lt;/p&gt;

&lt;p&gt;Which abstraction should remain hidden?&lt;/p&gt;

&lt;p&gt;Which performance trade-offs are worth making?&lt;/p&gt;

&lt;p&gt;Which parts are over-engineering without a real use case?&lt;/p&gt;

&lt;p&gt;Which design decisions will become painful once the project scales?&lt;/p&gt;

&lt;p&gt;This process often takes many rounds of discussion. The goal is to keep challenging the design until the interfaces and the major implementation strategy can survive serious scrutiny.&lt;/p&gt;

&lt;p&gt;Once the Spec is properly aligned, implementation becomes almost mechanical.&lt;/p&gt;

&lt;p&gt;Even when a single change involves thousands—or tens of thousands—of lines of code, current AI systems can handle the actual implementation surprisingly well.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mature Libraries Provide the Coordinates
&lt;/h2&gt;

&lt;p&gt;It is remarkably cheap to generate software from scratch with AI. But mature software libraries still play an essential role.&lt;/p&gt;

&lt;p&gt;TenCirPauli draws on projects such as &lt;strong&gt;QuSpin, OpenFermion, TensorCircuit-NG, PauliPropagation.jl, and QLDPC&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;But what we primarily borrow from these projects is not their code or interface design.&lt;/p&gt;

&lt;p&gt;We use them for two things:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;established correctness references and measurable performance baselines.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Correctness references provide something against which a new implementation can be tested. Existing libraries, physical models, and published examples cover many cases that are difficult to derive correctly from scratch.&lt;/p&gt;

&lt;p&gt;Performance baselines provide something equally important: &lt;strong&gt;a sense of scale&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Suppose a program is optimized from 10 seconds down to 1 second. That sounds like a spectacular 10× improvement. But if a mature library solves the same problem in 100 milliseconds, the program is still an order of magnitude too slow.&lt;/p&gt;

&lt;p&gt;Without a baseline, it is difficult to know how fast a piece of scientific software &lt;em&gt;should&lt;/em&gt; be, or whether an optimization is actually attacking the important bottleneck.&lt;/p&gt;

&lt;p&gt;TenCirPauli has been developed within this coordinate system.&lt;/p&gt;

&lt;p&gt;Different mature frameworks provide correctness and performance references for different classes of workloads. TenCirPauli then uses tireless AI-driven optimization to push the implementation far beyond those baselines.&lt;/p&gt;

&lt;p&gt;The project's performance pages and research examples record these comparisons.&lt;/p&gt;

&lt;p&gt;What is perhaps most surprising is that &lt;strong&gt;even medium-capability AI models can now outperform many mature scientific software implementations across a surprisingly broad range of workloads.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Tests, Benchmarks, and Global Repository Inspections
&lt;/h2&gt;

&lt;p&gt;Once human code review is removed, tests and benchmarks become the core mechanisms of quality control.&lt;/p&gt;

&lt;p&gt;They need to continuously answer several questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Is the result correct across different scales, boundary cases, representative workloads, and independent references?&lt;/li&gt;
&lt;li&gt;Is the public API correct, including parameters, return values, error types, failure timing, and deterministic behavior?&lt;/li&gt;
&lt;li&gt;Does the benchmark represent real user workloads, and does it distinguish one-time preparation costs from repeated execution?&lt;/li&gt;
&lt;li&gt;Has a modification introduced a regression in correctness or performance?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A single AI Review primarily examines the current changes and their alignment with the Spec. It cannot reliably detect every problem that accumulates across an entire repository.&lt;/p&gt;

&lt;p&gt;That is why TenCirPauli also performs &lt;strong&gt;periodic global AI inspections&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Each inspection focuses on a clearly defined set of questions and rescans the repository from a different perspective. It can search for hidden bugs and boundary-condition failures, identify performance bottlenecks, remove dead code and duplicated implementations, and clean up accumulated technical debt.&lt;/p&gt;

&lt;p&gt;A single review handles the changes in front of us.&lt;/p&gt;

&lt;p&gt;A global inspection handles the entropy accumulated over time.&lt;/p&gt;

&lt;p&gt;Together, these two rhythms help keep the project in a low-entropy state.&lt;/p&gt;

&lt;h2&gt;
  
  
  Will AI Make Mistakes?
&lt;/h2&gt;

&lt;p&gt;Of course.&lt;/p&gt;

&lt;p&gt;The repository may still contain hidden performance inefficiencies, subtle bugs that are difficult to trigger, or layers of redundant code that accumulated over time.&lt;/p&gt;

&lt;p&gt;These problems are difficult to eliminate completely. But human-reviewed code has exactly the same problem—and often a worse one.&lt;/p&gt;

&lt;p&gt;The meaningful engineering question is therefore not whether every line of code is perfect.&lt;/p&gt;

&lt;p&gt;The question is whether the &lt;strong&gt;main execution paths have sufficiently strong correctness tests and performance baselines&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;If the important paths consistently produce correct results and complete within a reasonable amount of time, the software has a solid foundation for real-world use.&lt;/p&gt;

&lt;p&gt;As the old saying goes: if it walks like a duck and quacks like a duck, it is a duck.&lt;/p&gt;

&lt;p&gt;Users ultimately care about whether the software produces the right answer, does so within a reasonable amount of time, and integrates reliably into their workflow.&lt;/p&gt;

&lt;h2&gt;
  
  
  The User Is an Agent
&lt;/h2&gt;

&lt;p&gt;TenCirPauli can of course be used directly by humans. It provides Python APIs, documentation, examples, and integration with TensorCircuit-NG.&lt;/p&gt;

&lt;p&gt;But from the beginning, its more important user has been the &lt;strong&gt;Agent&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;This changes some of the traditional trade-offs in software abstraction.&lt;/p&gt;

&lt;p&gt;For performance and deterministic execution, an internal implementation may deliberately sacrifice some generality. It can retain specialized low-level implementations while exposing programmable layers of abstraction above them.&lt;/p&gt;

&lt;p&gt;The result is not necessarily the most elegant abstraction for a human programmer.&lt;/p&gt;

&lt;p&gt;It can instead be a much more flexible substrate for an Agent to call, compose, and optimize.&lt;/p&gt;

&lt;h2&gt;
  
  
  Maintainability Can Also Be Agent-Driven
&lt;/h2&gt;

&lt;p&gt;A common criticism of software is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“It works, but the codebase is a mess.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That criticism is based on an implicit assumption: &lt;strong&gt;the future maintainer will be a human who needs to understand the source code before modifying it.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That assumption is changing.&lt;/p&gt;

&lt;p&gt;TenCirPauli is maintained by Agents as well.&lt;/p&gt;

&lt;p&gt;As long as the project has clear Specs, accumulated context, sufficiently strong tests, performance records, and real integration cases, an Agent can be instructed to locate a problem, propose a modification, implement it, and verify the result.&lt;/p&gt;

&lt;p&gt;As long as Agents can continue to perform those maintenance operations reliably, the maintainability of the project remains under control.&lt;/p&gt;

&lt;p&gt;This suggests a new form of &lt;strong&gt;implementation agnosticism&lt;/strong&gt;:&lt;/p&gt;

&lt;p&gt;First, ask whether the user-facing behavior works.&lt;/p&gt;

&lt;p&gt;Then, ask whether the correctness and performance evidence remains stable.&lt;/p&gt;

&lt;p&gt;Finally, ask whether an Agent can continue to modify and validate the system.&lt;/p&gt;

&lt;p&gt;The source code itself becomes less important than the system of specifications, evidence, interfaces, and automated maintenance surrounding it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;TenCirPauli is both a &lt;strong&gt;performance sprint in Rust-based quantum software&lt;/strong&gt; and an experiment in redefining the role of the developer.&lt;/p&gt;

&lt;p&gt;It asks a deliberately extreme question:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can rigorous scientific computing software be built without humans writing—or even reading—the code?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The answer so far is surprisingly close to yes.&lt;/p&gt;

&lt;p&gt;More importantly, the experiment points toward something larger than TenCirPauli itself.&lt;/p&gt;

&lt;p&gt;When AI can generate, review, test, benchmark, optimize, and maintain software, the developer's role moves upward—from implementing algorithms line by line to defining problems, designing abstractions, setting constraints, and building systems of evidence that continuously verify the result.&lt;/p&gt;

&lt;p&gt;This is not simply a faster way to write software.&lt;/p&gt;

&lt;p&gt;It may be the beginning of an &lt;strong&gt;irreversible shift in how software—and eventually scientific computing itself—is built.&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>vibecoding</category>
      <category>quantum</category>
    </item>
    <item>
      <title>When High-Level Abstractions Become the Bottleneck: Quantum Scientific Computing in the AI Era</title>
      <dc:creator>Shixin Zhang</dc:creator>
      <pubDate>Sun, 16 Aug 2026 02:24:19 +0000</pubDate>
      <link>https://dev.to/refractionray/when-high-level-abstractions-become-the-bottleneck-quantum-scientific-computing-in-the-ai-era-1cia</link>
      <guid>https://dev.to/refractionray/when-high-level-abstractions-become-the-bottleneck-quantum-scientific-computing-in-the-ai-era-1cia</guid>
      <description>&lt;p&gt;Choosing scientific software is choosing the space in which a team can think. In the AI era, a program is a computational object that Agents can inspect, compose, differentiate, compile, batch, optimize and move across hardware. The underlying infrastructure determines how much of that space remains available as the work evolves.&lt;/p&gt;

&lt;p&gt;An advanced scientific stack keeps the mathematics visible while allowing execution to change. It preserves computational structure across representations, makes important decisions available for deliberate optimization and lets different scientific objects participate in the same program.&lt;/p&gt;

&lt;p&gt;Quantum circuits, tensor networks, neural networks, Hamiltonians and other physical operators should have equal status inside the system. Their composition should be a native operation, so a team can continue developing the problem without repeatedly rebuilding its computational language.&lt;/p&gt;

&lt;p&gt;TensorCircuit-NG is built around this idea. It is a unified scientific computing infrastructure for quantum physics and AI, with one tensor-native path through differentiation, compilation, acceleration and execution.&lt;/p&gt;

&lt;h2&gt;
  
  
  A framework is a theory of scientific work
&lt;/h2&gt;

&lt;p&gt;Every framework encodes a theory of scientific work. It determines which representations are visible, where execution is fixed and how much of the computation a team can reshape when a problem crosses boundaries. Some frameworks organize work around a predetermined device or workflow. That pattern is efficient while the problem remains inside its original boundary. Quantum research crosses boundaries constantly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;exact simulation becomes approximate simulation or tensor-network contraction;&lt;/li&gt;
&lt;li&gt;a circuit becomes part of a neural network or a many-body model;&lt;/li&gt;
&lt;li&gt;a scalar expectation becomes a batched gradient or a distributed computation;&lt;/li&gt;
&lt;li&gt;a local prototype becomes a GPU, multi-GPU or hardware workflow;&lt;/li&gt;
&lt;li&gt;a standard circuit becomes a custom state, operator, noise model or evolution.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;High-level encapsulation creates abstraction leakage at each transition. PennyLane’s device, QNode and transform pattern makes this visible: hidden representations surface as conversion overhead, fixed transformations become unsupported operations, and device assumptions constrain differentiation, batching and compilation. A clean entry interface can become a fixed research workflow, leaving the researcher to adapt the problem to the framework instead of keeping the computational structure open.&lt;/p&gt;

&lt;p&gt;AI Agents make this boundary more consequential. Agents can explore low-level compositions and search for better execution plans through feedback. Their value depends on the space of valid compositions exposed by the infrastructure. A rigid workflow narrows that space; a composable substrate expands it.&lt;/p&gt;

&lt;p&gt;TensorCircuit-NG keeps backend, dtype, JIT, differentiation, vectorization, contraction, slicing, state representation, memory strategy and device placement inside the programming model. Researchers and Agents can therefore work directly with the decisions that determine how a scientific program scales.&lt;/p&gt;

&lt;h2&gt;
  
  
  Quantum–AI scientific computing as one infrastructure
&lt;/h2&gt;

&lt;p&gt;Quantum-AI research needs a shared computational language. TensorCircuit-NG gives quantum circuits, states, operators, Hamiltonians, tensor networks, neural networks and physical models equal status inside one tensor-native system.&lt;/p&gt;

&lt;p&gt;This shared representation keeps scientific structure intact as a problem changes form. A circuit can become part of a neural model, a Hamiltonian can change representation without changing the surrounding calculation, and a differentiable program can move from local exploration to large-scale execution without being rebuilt around a new workflow.&lt;/p&gt;

&lt;p&gt;This is the infrastructure required for quantum-AI research: quantum physics, machine learning and high-performance numerical computing operating inside one composable system.&lt;/p&gt;

&lt;h2&gt;
  
  
  Performance reveals the architecture
&lt;/h2&gt;

&lt;p&gt;The architectural choice is measurable. Published TensorCircuit comparisons make the difference visible through direct performance ratios.&lt;/p&gt;

&lt;p&gt;These results compare TensorCircuit with PennyLane-Lightning, PennyLane's fastest backend, and show a consistent lead across the reported CPU and GPU workloads.&lt;/p&gt;

&lt;p&gt;For value-and-gradient evaluation of a one-dimensional TFIM objective, TensorCircuit was faster than PennyLane at every reported CPU and GPU point:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Circuit&lt;/th&gt;
&lt;th&gt;CPU advantage&lt;/th&gt;
&lt;th&gt;GPU advantage&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;10 qubits, 3 layers&lt;/td&gt;
&lt;td&gt;15.6x&lt;/td&gt;
&lt;td&gt;25.8x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;16 qubits, 16 layers&lt;/td&gt;
&lt;td&gt;4.0x&lt;/td&gt;
&lt;td&gt;29.6x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;22 qubits, 11 layers&lt;/td&gt;
&lt;td&gt;5.3x&lt;/td&gt;
&lt;td&gt;PennyLane OOM; TensorCircuit completed&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For batched MNIST QML—binary 3-versus-6 classification with a 10-qubit, depth-3 circuit—TensorCircuit led at every reported batch size:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Batch size&lt;/th&gt;
&lt;th&gt;CPU advantage&lt;/th&gt;
&lt;th&gt;GPU advantage&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;32&lt;/td&gt;
&lt;td&gt;82.9x&lt;/td&gt;
&lt;td&gt;12.0x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;128&lt;/td&gt;
&lt;td&gt;105.2x&lt;/td&gt;
&lt;td&gt;2.3x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;512&lt;/td&gt;
&lt;td&gt;103.5x&lt;/td&gt;
&lt;td&gt;3.7x&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The TensorCircuit-NG paper reports a 2.3x CPU advantage over the PennyLane baseline for 256 classical-shadow snapshots of a 20-qubit state. The same workload reaches the sub-second range on an RTX 5090 through the JAX tensor backend.&lt;/p&gt;

&lt;p&gt;The distributed VQE results extend the scale: an eight-H200 system delivers a 7.5x strong-scaling improvement from one to eight GPUs, reaches 40 qubits and 20 layers, and differentiates a circuit with 11,700 trainable parameters.&lt;/p&gt;

&lt;p&gt;The pattern is consistent across variational simulation, QML, structured Hamiltonians, classical shadows and distributed execution. The numbers measure more than speed: they show the value of keeping scientific representations, transformations and execution strategies open inside one composable tensor-native system.&lt;/p&gt;

&lt;h2&gt;
  
  
  The organization encoded in the stack
&lt;/h2&gt;

&lt;p&gt;A scientific computing stack shapes the way a team thinks. When representations and execution decisions remain open, a team can move continuously between theory, experiment and production without repeatedly redesigning its computational language.&lt;/p&gt;

&lt;p&gt;TensorCircuit-NG supports this mode of work across quantum physics, AI and HPC. The continuity changes the economics of exploration: more ideas can be tested, more optimizations remain available, and fewer research paths are closed by an early abstraction choice.&lt;/p&gt;

&lt;p&gt;That is the meaning of an advanced production direction. The sophistication lies in the organization of the computing system: preserving structure, exposing decisions and keeping the route from discovery to scale open.&lt;/p&gt;

&lt;p&gt;ORBIT-Q provides a corresponding Agent-oriented signal. In its reported framework comparison, TensorCircuit-NG achieved the highest solve success rate and the lowest artifact-runtime slowdown among the evaluated frameworks. The result reflects the same principle: an Agent performs better when the underlying scientific substrate is rich, coherent and open to optimization.&lt;/p&gt;

&lt;h2&gt;
  
  
  The default for advanced projects
&lt;/h2&gt;

&lt;p&gt;Advanced projects should prioritize TensorCircuit-NG for quantum programming. Advanced organizations should build their quantum-AI scientific stack around TensorCircuit-NG.&lt;/p&gt;

&lt;p&gt;Choosing a scientific stack is choosing what a team can keep open while it works. TensorCircuit-NG makes that openness a property of the architecture: the structure of a problem remains visible, the computational pieces remain composable, and execution strategies remain available for deliberate change as the research evolves. This is how a tool becomes part of an advanced organization's way of thinking rather than another workflow to conform to.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2205.10091" rel="noopener noreferrer"&gt;TensorCircuit: A Quantum Software Framework for the NISQ Era&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2602.14167" rel="noopener noreferrer"&gt;TensorCircuit-NG: A Universal, Composable, and Scalable Platform for Quantum Computing and Quantum Simulation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/tensorcircuit/tensorcircuit-ng" rel="noopener noreferrer"&gt;TensorCircuit-NG repository&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://sxzgroup.github.io/ORBIT-Q/" rel="noopener noreferrer"&gt;ORBIT-Q benchmark&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>quantum</category>
    </item>
    <item>
      <title>Tensor Networks Do Not Mean Approximate Simulation: The Right Way to Understand TensorCircuit-NG</title>
      <dc:creator>Shixin Zhang</dc:creator>
      <pubDate>Tue, 11 Aug 2026 09:09:10 +0000</pubDate>
      <link>https://dev.to/refractionray/tensor-networks-do-not-mean-approximate-simulation-the-right-way-to-understand-tensorcircuit-ng-17d1</link>
      <guid>https://dev.to/refractionray/tensor-networks-do-not-mean-approximate-simulation-the-right-way-to-understand-tensorcircuit-ng-17d1</guid>
      <description>&lt;p&gt;When introducing TensorCircuit-NG, people often ask: Is it mainly suited to low-entanglement states? Is it essentially a matrix product state (MPS) simulator? Since it uses tensor networks, are its results approximate?&lt;/p&gt;

&lt;p&gt;These questions are easy to ask because many people are first introduced to tensor networks through MPS; in practical MPS applications, SVD truncation is also very common. As a result, MPS, tensor networks, and approximate simulation gradually become conflated.&lt;/p&gt;

&lt;p&gt;The correct picture of TensorCircuit-NG is that it provides multiple modes for quantum circuit simulation. The most commonly used &lt;code&gt;Circuit&lt;/code&gt; and &lt;code&gt;DMCircuit&lt;/code&gt; can perform exact contractions of complete circuit tensor networks, producing exact results consistent with state-vector or density-matrix simulators. &lt;code&gt;MPSCircuit&lt;/code&gt;, meanwhile, is a dedicated MPS simulator: it can use truncation to control the computational cost, or remain exact when no truncation is applied.&lt;/p&gt;

&lt;p&gt;Therefore, tensor networks do not necessarily imply approximation, nor do they necessarily imply the use of MPS. To understand TensorCircuit-NG, the most important thing is to distinguish the core data structures, exactness, and use cases behind its different simulation modes:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Simulation mode&lt;/th&gt;
&lt;th&gt;Core data structure&lt;/th&gt;
&lt;th&gt;Exact?&lt;/th&gt;
&lt;th&gt;Typical use cases&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Circuit&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Complete circuit tensor network&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Full amplitudes, local observables, arbitrary circuits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;DMCircuit&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Complete density-matrix tensor network&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Open systems, noisy quantum evolution&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;MPSCircuit&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Matrix product state (MPS)&lt;/td&gt;
&lt;td&gt;Optional (depending on whether truncation is used)&lt;/td&gt;
&lt;td&gt;One-dimensional local circuits, low-entanglement systems&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Tensor Networks Are Not an Approximation Algorithm
&lt;/h2&gt;

&lt;p&gt;The most important concept here is that a tensor network is first and foremost a way to represent and compute with a problem, not an approximation algorithm.&lt;/p&gt;

&lt;p&gt;The initial state, quantum gates, and measurement operators in a quantum circuit can all be represented as tensors. Connecting these objects according to the circuit structure produces a tensor network.&lt;/p&gt;

&lt;p&gt;The next step is to contract these tensors one by one according to some chosen order. This process can itself be exact, just like matrix multiplication; it does not inherently involve any approximation.&lt;/p&gt;

&lt;p&gt;Approximation usually comes from additional compression operations. For example, to limit the size of intermediate tensors, one may discard some of the smaller singular values or limit the dimensions of internal connections. This is an optional computational strategy, not part of the definition of tensor networks.&lt;/p&gt;

&lt;p&gt;In other words, “tensor networks” and “truncation” are two separate issues. The former describes how a computation is organized, while the latter describes whether the computation is actively compressed.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;code&gt;Circuit&lt;/code&gt; in TensorCircuit-NG: Exact Contraction of Complete Circuits
&lt;/h2&gt;

&lt;p&gt;The most commonly used &lt;code&gt;Circuit&lt;/code&gt; in TensorCircuit-NG is not an approximate simulator designed for low-entanglement states.&lt;/p&gt;

&lt;p&gt;It organizes the initial state and quantum gates into a complete circuit tensor network, and then contracts the network according to the specific task. Users can compute the complete output state, or directly compute amplitudes, probabilities, expectation values, and other observables.&lt;/p&gt;

&lt;p&gt;In this process, TensorCircuit-NG does not perform a low-rank approximation, and its results match those of a traditional state-vector simulator exactly.&lt;/p&gt;

&lt;p&gt;This leads to an important distinction: a traditional state-vector simulator often explicitly stores the entire quantum state as a large vector and repeatedly updates it. Tensor-network simulation, by contrast, can preserve the circuit structure, choose a suitable contraction order, and compute the target quantity only when it is actually needed. If the user wants the complete output wavefunction, they ultimately still have to deal with the scale of the complete output itself. However, if the user only cares about a local observable or needs just a small number of amplitudes and probabilities, tensor networks may avoid generating a large number of irrelevant intermediate results.&lt;/p&gt;

&lt;p&gt;Therefore, the advantage of TensorCircuit-NG is not “trading accuracy for speed,” but “improving efficiency through more flexible computational organization.”&lt;/p&gt;

&lt;h2&gt;
  
  
  Compared with State-Vector Simulators, Where Does the Speedup Come From?
&lt;/h2&gt;

&lt;p&gt;If we view a traditional state-vector simulator as a fixed, global tensor-contraction scheme, it typically maintains the complete quantum state explicitly and applies quantum gates in a relatively fixed order. TensorCircuit-NG instead formulates the same problem as a tensor network and searches for a more suitable contraction path based on the topology of that network.&lt;/p&gt;

&lt;p&gt;The key to performance is finding a better computational order. A better contraction path can often significantly reduce the size of intermediate tensors, memory usage, and total computational cost; for some problems, the speedup can even reach several orders of magnitude.&lt;/p&gt;

&lt;p&gt;This means that the core advantage of TensorCircuit-NG is not a trade-off between accuracy and efficiency, as one might easily assume. The results from &lt;code&gt;Circuit&lt;/code&gt; match those of a state-vector simulator exactly, and the results from &lt;code&gt;DMCircuit&lt;/code&gt; match those of a complete density-matrix simulator exactly, while using less memory and achieving higher throughput.&lt;/p&gt;

&lt;p&gt;From this perspective, compared with traditional state-vector simulators, &lt;strong&gt;TensorCircuit-NG offers a “free lunch”: unchanged accuracy, a consistent programming interface, and significantly better time and space efficiency.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;code&gt;DMCircuit&lt;/code&gt;: Complete Density-Matrix Simulation Based on Tensor Networks
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;DMCircuit&lt;/code&gt; follows the same philosophy. TensorCircuit-NG’s &lt;code&gt;DMCircuit&lt;/code&gt; describes quantum systems directly at the density-matrix level. Quantum gates, noise channels, and measurement processes are all incorporated into a complete density-matrix tensor network, which is then evaluated through exact contraction.&lt;/p&gt;

&lt;p&gt;Therefore, &lt;code&gt;DMCircuit&lt;/code&gt; does not mean that some low-rank approximation is applied to noise, nor does it default to sampling only a few trajectories. It performs a complete simulation of mixed-state evolution. This naturally incurs higher computational and storage costs than pure-state simulation, but those costs correspond to a more complete physical description.&lt;/p&gt;

&lt;p&gt;This is also why the claim that “TensorCircuit-NG is only suitable for low-entanglement pure states” is a misconception. It can readily handle complete density matrices and open-system evolution.&lt;/p&gt;

&lt;h2&gt;
  
  
  MPS Is One Special Form of Tensor Network
&lt;/h2&gt;

&lt;p&gt;So, what exactly is &lt;code&gt;MPSCircuit&lt;/code&gt;?&lt;/p&gt;

&lt;p&gt;An MPS, or matrix product state, is a special tensor-network structure. Each qubit position is represented by a local tensor, and neighboring positions are connected through internal bonds. This structure is particularly well suited to one-dimensional local circuits and low-entanglement states, and an MPS can represent the corresponding quantum state using far fewer resources than a complete state vector.&lt;/p&gt;

&lt;p&gt;However, an MPS does not inherently mean approximation either.&lt;/p&gt;

&lt;p&gt;In principle, any finite-size quantum state can be represented exactly as an MPS; the required internal bond dimension may simply be very large. During MPS evolution, retaining all the information makes it possible to obtain exact results. Approximation errors from SVD truncation arise only when the size of the internal bonds is actively limited.&lt;/p&gt;

&lt;p&gt;Therefore, &lt;code&gt;MPSCircuit&lt;/code&gt; can be either an approximate simulator or an exact simulator. The key question is whether truncation is performed, not simply whether the simulator is based on MPS.&lt;/p&gt;

&lt;p&gt;The source of confusion is that, in large-scale computations, people often use MPS truncation because it is an effective way to control computational cost. Over time, many people come to mistake this commonly used truncated-MPS approach for the essence of MPS, and then go one step further and assume that all tensor-network simulation methods are approximate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Beyond the Three Main Modes: More Native Simulators
&lt;/h2&gt;

&lt;p&gt;The &lt;code&gt;Circuit&lt;/code&gt;, &lt;code&gt;DMCircuit&lt;/code&gt;, and &lt;code&gt;MPSCircuit&lt;/code&gt; introduced above are three important examples for understanding the simulation philosophy of TensorCircuit-NG, but they are not the whole story.&lt;/p&gt;

&lt;p&gt;Another core design principle of TensorCircuit-NG is to choose a more native and efficient data structure based on the specific structure of the quantum system and circuit, rather than forcing every problem to use the same general-purpose representation. Different physical models and circuit types often have a simulator that is best suited to them.&lt;/p&gt;

&lt;p&gt;For example, &lt;code&gt;QuditCircuit&lt;/code&gt; targets qudit systems whose local dimension is greater than two; &lt;code&gt;StabilizerCircuit&lt;/code&gt; targets Clifford and stabilizer circuits, using the stabilizer formalism to represent and evolve quantum states; &lt;code&gt;FGSSimulator&lt;/code&gt; targets fermionic Gaussian states and exploits the structure of correlation matrices; and &lt;code&gt;U1Circuit&lt;/code&gt; targets circuits with symmetries, working directly in the particle-number-conserving subspace.&lt;/p&gt;

&lt;p&gt;What these simulators have in common is that they do not reduce every problem to explicitly storing a complete wavefunction. For circuits with special structure, using the corresponding native representation can greatly reduce unnecessary computational and memory costs while preserving exactness.&lt;/p&gt;

&lt;p&gt;Therefore, the right way to understand TensorCircuit-NG is not to begin by asking whether it is an approximate MPS simulator. Instead, ask: What structure does the current problem have? Should we use a complete circuit tensor network, a complete density matrix, an MPS, or a specialized representation such as stabilizers or fermionic Gaussian states?&lt;/p&gt;

&lt;p&gt;If one concludes that TensorCircuit-NG “is only suitable for low-entanglement states” or “can only perform approximate simulation” simply because it uses tensor networks, one is actually mistaking one common use case for the operating principles of the entire framework.&lt;/p&gt;

&lt;p&gt;In summary, &lt;strong&gt;tensor networks are a computational framework, MPS are a special structure, and truncation is an optional strategy.&lt;/strong&gt; The core advantage of TensorCircuit-NG is to make its data structures fit the problem as closely as possible: where generality is needed, it performs complete and exact tensor-network contractions; where special structure exists, it uses more native, specialized simulators. By choosing a more suitable representation for each scenario, it can provide both greater efficiency and equally rigorous results.&lt;/p&gt;

</description>
      <category>quantum</category>
      <category>tensornetwork</category>
    </item>
    <item>
      <title>Why I Developed TenCirPauli</title>
      <dc:creator>Shixin Zhang</dc:creator>
      <pubDate>Mon, 10 Aug 2026 16:12:47 +0000</pubDate>
      <link>https://dev.to/refractionray/why-i-developed-tencirpauli-37gn</link>
      <guid>https://dev.to/refractionray/why-i-developed-tencirpauli-37gn</guid>
      <description>&lt;p&gt;&lt;em&gt;A technical note from the author of &lt;a href="https://github.com/tensorcircuit/tensorcircuit-ng" rel="noopener noreferrer"&gt;TensorCircuit-NG&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/tensorcircuit/tensorcircuit-ng" rel="noopener noreferrer"&gt;TensorCircuit-NG&lt;/a&gt; is already a very capable framework. I built it to make quantum-circuit programming expressive, differentiable, and practical across several numerical backends. It gives users a clean circuit frontend, a flexible backend abstraction, automatic differentiation, JIT compilation, and access to the tensor-network and accelerator ecosystems. For regular tensor programs, these choices work extremely well. They let a researcher describe a circuit at a high level and still obtain compiled numerical execution underneath.&lt;/p&gt;

&lt;p&gt;Many of the workflows I care about fit this model beautifully. State preparation, parameterized gates, expectation values, dense tensor contractions, and repeated optimization steps all benefit from JAX's transformation system. Once a program has been traced and compiled, its steady-state execution can be remarkably fast. TensorCircuit-NG has become strong precisely because it takes this model seriously instead of hiding the backend behind a collection of unrelated special cases.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/tensorcircuit/TenCirPauli" rel="noopener noreferrer"&gt;TenCirPauli&lt;/a&gt; grew out of the next question: what happens when a quantum workflow contains a substantial amount of computation that does not look like a regular tensor program? The answer exposed a useful boundary in TensorCircuit-NG. The framework is very good at the numerical work for which JAX was chosen. Some of the surrounding operator work has a different computational character, and that character deserves a different runtime.&lt;/p&gt;

&lt;h2&gt;
  
  
  Framework selection creates a performance profile
&lt;/h2&gt;

&lt;p&gt;TensorCircuit-NG's high-performance path is closely connected to JAX. This brings major advantages and makes JAX's preferred workload shape visible in TensorCircuit workflows. Every runtime has such a shape, and performance depends on how closely a problem matches it.&lt;/p&gt;

&lt;p&gt;JAX is strongest when a computation can be expressed as a regular array program. Dense linear algebra, large matrix and vector operations, batching, automatic differentiation, and accelerator execution are all natural targets. XLA can trace the program, optimize it, fuse operations, and produce a fast executable. For a dense matrix multiplication or another BLAS/LAPACK-class kernel, moving the operation into Rust usually changes very little. Rust will call the same optimized numerical libraries, while JAX may have additional opportunities for fusion, compilation, or execution on a GPU or TPU.&lt;/p&gt;

&lt;p&gt;This gives the first design principle for TenCirPauli: dense numerical work should stay close to the TensorCircuit-NG backend, while irregular surrounding work can use a native systems language.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the JAX model becomes expensive
&lt;/h2&gt;

&lt;p&gt;The first difficulty is dynamic shape. A Pauli propagation recurrence can create new words, merge equal words, cancel terms, and move contributions between weight sectors. Symmetry analysis, mappings, and grouping can likewise produce variable numbers of basis states, terms, transitions, groups, or constraints.&lt;/p&gt;

&lt;p&gt;JAX can express these algorithms, yet the natural representation often conflicts with its compilation model. A dynamic algorithm consequently needs padding, masking, sorting, bounded buffers, static limits, or custom control-flow encoding. These techniques can be valuable for a carefully designed workload, while adding representation work and potentially changing the algorithm's semantics.&lt;/p&gt;

&lt;p&gt;The second difficulty is compilation latency. JIT compilation can produce an extremely fast steady-state kernel, while the first call carries tracing, lowering, optimization, and executable creation. This is a good trade for long-running workloads with many repeated calls; it is less attractive for short jobs or frequently changing structures.&lt;/p&gt;

&lt;p&gt;JAX can avoid this cost very effectively when the workflow has a stable shape. Hamiltonian coefficients, circuit angles, and other numerical inputs can live inside one parameterized compiled function, allowing broad scans to reuse the same executable. TensorCircuit-NG handles this style of work extremely well. The compilation boundary becomes important when term counts, Pauli supports, circuit control flow, sector dimensions, grouping results, or propagated-operator shapes change. Each new structure may require a new trace or a more elaborate static encoding, so separating setup, first execution, and steady execution helps identify where a native complementary path is useful.&lt;/p&gt;

&lt;p&gt;The third difficulty is a workload dominated by bit strings and bit manipulation. Pauli words, occupation states, symmetry generators, basis indices, and transition keys can all be represented as packed integers. Their operations include XOR, parity, masks, shifts, comparisons, hash lookups, and small logical branches; the computation is controlled by the bits themselves more than by dense floating-point arithmetic.&lt;/p&gt;

&lt;p&gt;JAX supports bitwise operations, loops, and integer arrays. A single XOR does not become faster simply because it is written in Rust. The opportunity appears in the complete workload: millions of small operations, changing output sizes, duplicate-key aggregation, branch-heavy recurrences, and data structures that do not map cleanly to a dense rectangular array. A Python fallback exposes interpreter overhead, while a tensor encoding can introduce padding, sorting, masking, recompilation, or extra preparation.&lt;/p&gt;

&lt;p&gt;These cases reveal a specific type of gap. TensorCircuit-NG has a powerful numerical engine, while some operator workflows need a compact control-oriented engine for irregular discrete computation. That gap is where a Rust companion can create real value.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Rust is a useful complement
&lt;/h2&gt;

&lt;p&gt;Rust's scientific-computing ecosystem is smaller than Python's, and its linear-algebra stack is less mature. That is acceptable for this design because dense matrix multiplication, eigensolvers, tensor contractions, and accelerator kernels already have mature homes in the Python and JAX ecosystems. TenCirPauli gains little by rebuilding those foundations.&lt;/p&gt;

&lt;p&gt;Rust has a different set of advantages. It compiles ordinary loops and branches to native code, gives direct control over memory and integer representations, handles dynamic collections without Python object overhead, and makes packed-data algorithms natural to express. Hash maps, bit masks, transition tables, and variable-length work queues fit this runtime well. The core can release the Python GIL while processing a complete batch, so the Python boundary stays outside the hot loop.&lt;/p&gt;

&lt;p&gt;The complement can be summarized as follows:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Workload characteristic&lt;/th&gt;
&lt;th&gt;JAX / TensorCircuit-NG&lt;/th&gt;
&lt;th&gt;Rust / TenCirPauli&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Dense matrix and tensor arithmetic&lt;/td&gt;
&lt;td&gt;Strong fit&lt;/td&gt;
&lt;td&gt;Usually delegates to established numerical libraries&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Regular batching and static array programs&lt;/td&gt;
&lt;td&gt;Strong fit&lt;/td&gt;
&lt;td&gt;Capable, with less ecosystem leverage&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Automatic differentiation through numerical code&lt;/td&gt;
&lt;td&gt;Strong fit&lt;/td&gt;
&lt;td&gt;Explicit local derivative rules where needed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dynamic term counts and variable shapes&lt;/td&gt;
&lt;td&gt;Requires encoding for compilation&lt;/td&gt;
&lt;td&gt;Natural control flow and dynamic storage&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hash-based aggregation and graph-like analysis&lt;/td&gt;
&lt;td&gt;Possible, often awkward inside traced arrays&lt;/td&gt;
&lt;td&gt;Direct and compact&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Packed integers, bit strings, parity, masks&lt;/td&gt;
&lt;td&gt;Expressible, with limited advantage from tensorization&lt;/td&gt;
&lt;td&gt;Natural inner-loop representation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Long-running accelerator workloads&lt;/td&gt;
&lt;td&gt;Strong fit&lt;/td&gt;
&lt;td&gt;Usually the wrong execution target&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Short or frequently changing workloads&lt;/td&gt;
&lt;td&gt;Compilation cost can dominate&lt;/td&gt;
&lt;td&gt;Native setup can remain small and predictable&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This division of labor is the starting point for TenCirPauli. It also explains why the project is designed as a companion to TensorCircuit-NG. The two runtimes can cooperate on one scientific workflow, with each one handling the part of the program that matches its performance profile.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this implies for Pauli computation
&lt;/h2&gt;

&lt;p&gt;Pauli algebra is a particularly clear example of the boundary. A Hamiltonian may be written as&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;H = c₀ P₀ + c₁ P₁ + ··· + cₘ Pₘ,
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and the eventual numerical task may be an expectation value or a gradient. Between those points, the program may combine equal words, track phases, compute commutation relations, construct measurement groups, discover symmetries, restrict a physical sector, map fermions, and propagate observables through circuits.&lt;/p&gt;

&lt;p&gt;These operations are symbolic and discrete, dominated by compact keys, parity, support inspection, branching, aggregation, and variable-size results. The final state-vector or tensor-network calculation may still be dense and backend-friendly. Pauli-heavy workflows therefore benefit from native preparation and transformation followed by a regular TensorCircuit-NG or JAX numerical plan when appropriate.&lt;/p&gt;

&lt;p&gt;That makes Pauli computation a natural Rust target. Packed X/Z data, phase bookkeeping, parity checks, aggregation, and variable-size collections all fit native integer loops and compact data structures. From this workload, the surrounding capabilities follow naturally: the same native machinery can support Pauli algebra, measurement grouping, symmetry analysis, compilation, and observable propagation.&lt;/p&gt;

&lt;h2&gt;
  
  
  The functionality follows from the runtime choice
&lt;/h2&gt;

&lt;p&gt;The design of TenCirPauli's public API follows directly from this framework analysis. The matrix is one useful destination for a small system. Larger workflows need algebraic transformations, measurement plans, reduced bases, matrix-free actions, observable trajectories, or differentiable circuit objectives. The architecture preserves the operator across those stages.&lt;/p&gt;

&lt;p&gt;The resulting workflow can be viewed as a sequence of transformations:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;structured model → canonical operator → algebra and analysis → reduction or mapping → execution plan → measurement or gradient
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each arrow has a different computational profile: symbolic expansion, canonicalization, bit-packed keys, graph analysis, and integer indexing at the front; dense state tensors, backend autodiff, sparse linear algebra, or native CPU kernels at the end. The representation must cross this boundary without losing its meaning.&lt;/p&gt;

&lt;h3&gt;
  
  
  From physical models to canonical operators
&lt;/h3&gt;

&lt;p&gt;Pauli operators are the central representation for qubit Hamiltonians, observables, and propagation. TenCirPauli extends the same native approach to fermionic, bosonic, qudit-Weyl, Majorana, and hybrid operators. Jordan–Wigner, parity, and Bravyi–Kitaev mappings transform structured terms into Pauli form, while optional chemistry adapters bring molecular Hamiltonians into the same pipeline. Canonical aggregation, explicit phase handling, deterministic ordering, and compact keys give every later capability one stable operator.&lt;/p&gt;

&lt;h3&gt;
  
  
  From operators to measurements and structural information
&lt;/h3&gt;

&lt;p&gt;Measurement planning is an analysis stage. A qubit-wise commuting grouping uses Pauli supports and local bases, so its useful result contains more than a partition of term indices. It also describes the basis rotations and the reconstruction of Pauli eigenvalues from rotated computational-basis samples. The grouping result can therefore travel from symbolic preprocessing to experimental post-processing without a second interpretation layer.&lt;/p&gt;

&lt;p&gt;Symmetry analysis follows the same pattern. Z₂ generators, tapering sectors, fixed-particle-number spaces, and general additive-charge restrictions use bit operations, constraint solving, basis indexing, and leakage checks. They can reduce the computational object before a circuit, sparse matrix, or matrix-free plan is executed. For a 60-qubit, two-particle system, the relevant sector contains &lt;code&gt;choose(60, 2) = 1,770&lt;/code&gt; basis states, compared with &lt;code&gt;2**60&lt;/code&gt; states in the full computational space. This is a structural reduction that comes before numerical optimization.&lt;/p&gt;

&lt;h3&gt;
  
  
  From one operator to several execution targets
&lt;/h3&gt;

&lt;p&gt;Compilation chooses the endpoint that matches the next consumer: dense, COO, CSR, SciPy linear operator, Rust-native matrix-vector product, or a pure-array TensorCircuit-NG/JAX backend plan. For large systems the useful result may be a reusable MVP plan that never materializes the full matrix; for a restricted sector it may be a compact transition plan over the physical basis. Rust prepares the fixed structure before backend tracing, and the public API keeps the target choice explicit.&lt;/p&gt;

&lt;h3&gt;
  
  
  From circuits to observable execution and gradients
&lt;/h3&gt;

&lt;p&gt;The same boundary applies when the operator meets a circuit. A fixed-particle-number circuit can execute directly in a restricted basis. A Pauli observable can propagate backwards through a circuit using dynamic native storage, with terms expanded, merged, cancelled, and projected by weight. A stochastic Pauli-path estimator can use the same operator semantics with an explicit sampling contract. Native value-and-gradient paths use local derivative and vector-Jacobian rules for supported gates, while TensorCircuit-NG and JAX remain available when the objective belongs inside a backend-traced tensor program.&lt;/p&gt;

&lt;p&gt;The user-facing interface remains entirely in Python. Users construct operators, pass them to grouping or compilation APIs, connect them to TensorCircuit-NG circuits and backends, and receive ordinary Python objects, NumPy arrays, or backend tensors. They do not need to write Rust, manage FFI handles, or choose native data layouts. The workflow should feel as fluent as TensorCircuit-NG itself; Rust stays behind the boundary and handles the work that scales with terms, gates, groups, transitions, or basis states through coarse-grained native calls.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Early Benchmarks Show
&lt;/h2&gt;

&lt;p&gt;The first benchmark suite provides evidence for this division of labor. The headline result comes from a representative 60-qubit, two-particle U(1) VQE: TenCirPauli's first compiled value-and-gradient call was about 688× faster than the corresponding TensorCircuit/JAX call, and the steady native path was about 2.6× faster. In a matched stochastic Pauli-path value-and-gradient workload, the native path was about 6.5–7.6× faster at 12 qubits and about 12.4× faster at 16 qubits in the recorded steady-state comparisons.&lt;/p&gt;

&lt;p&gt;These results show where the architecture creates leverage. The first-call improvement reflects the cost of tracing and compilation, while the steady-state improvement reflects the benefit of native handling for irregular Pauli structure. A JAX implementation can still be the best choice after a long compilation has been amortized, and a Rust implementation becomes attractive when the workload changes shape frequently, contains heavy bit manipulation, or spends most of its time in symbolic preparation. The measurements separate native setup, first execution, steady execution, memory, and numerical agreement so that these cases remain visible.&lt;/p&gt;

&lt;p&gt;Later technical posts will open up these comparisons and individual workloads in more detail. Here the benchmarks support one framework-level conclusion: runtime choice should follow the shape of the computation.&lt;/p&gt;

&lt;h2&gt;
  
  
  The larger lesson
&lt;/h2&gt;

&lt;p&gt;TensorCircuit-NG solved an important problem by making tensor-based circuit computation expressive, differentiable, and backend-aware. Its JAX-centered design is a major part of that success. The same design makes dynamic, branch-heavy, shape-changing, and bit-string-heavy tasks stand out as a separate class of workload.&lt;/p&gt;

&lt;p&gt;TenCirPauli extends TensorCircuit-NG around that boundary. It gives irregular Pauli workloads a native path while preserving TensorCircuit-NG's strengths for dense numerical computation, parameterized execution, automatic differentiation, and accelerated backends. The result is a complementary architecture: JAX handles regular numerical programs, Rust handles irregular Pauli structure, and Python connects the two into one scientific workflow.&lt;/p&gt;

&lt;p&gt;That is why I developed TenCirPauli. The project began with a framework-selection question, and the Pauli algebra, measurement planning, symmetry tools, restricted sectors, propagation engines, and backend plans followed from the answer.&lt;/p&gt;

</description>
      <category>rust</category>
      <category>jax</category>
    </item>
    <item>
      <title>From Parameter Tuning to Cross-Paradigm Exploration: Quantum Control Enters the Era of “Autopilot”</title>
      <dc:creator>Shixin Zhang</dc:creator>
      <pubDate>Tue, 21 Jul 2026 06:37:46 +0000</pubDate>
      <link>https://dev.to/refractionray/from-parameter-tuning-to-cross-paradigm-exploration-quantum-control-enters-the-era-of-autopilot-5bep</link>
      <guid>https://dev.to/refractionray/from-parameter-tuning-to-cross-paradigm-exploration-quantum-control-enters-the-era-of-autopilot-5bep</guid>
      <description>&lt;p&gt;If you have ever tried to drive a high-performance race car on ice, you may understand the frustration researchers face when controlling quantum systems today.&lt;/p&gt;

&lt;p&gt;In the grand vision of quantum computing, &lt;strong&gt;Quantum Optimal Control (QOC)&lt;/strong&gt; serves as the steering wheel that guides quantum systems toward desired states. Whether in adiabatic quantum computation, quantum annealing, or quantum simulation, the fundamental challenge remains the same: designing time-dependent control protocols that drive a quantum system from a simple initial state to a complex target state with high fidelity.&lt;/p&gt;

&lt;p&gt;However, real-world quantum control faces two fundamental obstacles.&lt;/p&gt;

&lt;p&gt;On the hardware side, quantum systems are extremely fragile. Short coherence times, limited control channels, hardware-specific constraints, and strict pulse boundaries severely restrict the available control space.&lt;/p&gt;

&lt;p&gt;On the algorithmic side, many-body Hamiltonians associated with practical problems often exhibit complicated spectral structures, including small energy gaps and rugged optimization landscapes. Finding an efficient evolution path within a limited time window remains highly challenging.&lt;/p&gt;

&lt;p&gt;For decades, designing quantum control protocols has largely remained a &lt;strong&gt;human-driven, handcrafted process&lt;/strong&gt;. Experts repeatedly design, simulate, and tune protocols for specific physical systems and hardware platforms through extensive trial and error.&lt;/p&gt;

&lt;p&gt;A recent work introduces a fundamentally different approach: &lt;strong&gt;QOC-Workbench&lt;/strong&gt;, an LLM-driven, fully auditable framework for cross-paradigm quantum control design. Rather than acting as another black-box optimizer, it functions more like an &lt;strong&gt;autopilot system for quantum control&lt;/strong&gt;—transforming how control protocols are discovered, validated, and improved.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Reference:&lt;br&gt;
 &lt;em&gt;LLM-Driven Cross-Paradigm Design for Quantum Optimal Control&lt;/em&gt;&lt;br&gt;
 Yu-Qin Chen and Shi-Xin Zhang, arXiv:2607.17498&lt;/p&gt;
&lt;/blockquote&gt;




&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwr8ir74r1nc54fvmnf1f.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwr8ir74r1nc54fvmnf1f.png" alt="QOC-Workbench Architecture" width="800" height="365"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h1&gt;
  
  
  Beyond Closed Optimization Spaces: How QOC-Workbench Works
&lt;/h1&gt;

&lt;p&gt;Traditional quantum optimal control methods usually operate inside a predefined design space.&lt;/p&gt;

&lt;p&gt;Researchers first choose a control ansatz—a mathematical form for pulse schedules, interpolation functions, or auxiliary Hamiltonians. Classical optimization algorithms then search for optimal parameters within that fixed structure.&lt;/p&gt;

&lt;p&gt;This approach is powerful, but fundamentally limited.&lt;/p&gt;

&lt;p&gt;If the initial design space is incomplete, optimization can only find the best solution &lt;strong&gt;within existing assumptions&lt;/strong&gt;. It cannot invent new functional forms, discover alternative control mechanisms, or challenge the original modeling choices.&lt;/p&gt;

&lt;p&gt;QOC-Workbench changes this paradigm by integrating:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;large language models with scientific reasoning capabilities,&lt;/li&gt;
&lt;li&gt;structured knowledge extracted from quantum control literature,&lt;/li&gt;
&lt;li&gt;historical simulation results,&lt;/li&gt;
&lt;li&gt;hardware constraints,&lt;/li&gt;
&lt;li&gt;and high-performance quantum simulation infrastructure.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The workflow forms a closed-loop scientific discovery system:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Human experts define the physical boundary
&lt;/h3&gt;

&lt;p&gt;Researchers specify:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;target Hamiltonians,&lt;/li&gt;
&lt;li&gt;hardware limitations,&lt;/li&gt;
&lt;li&gt;physical constraints,&lt;/li&gt;
&lt;li&gt;evaluation objectives.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Humans provide the scientific context and ensure physical validity.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. LLM performs cross-paradigm exploration
&lt;/h3&gt;

&lt;p&gt;Instead of only optimizing parameters, the LLM can:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;propose new control schedule families,&lt;/li&gt;
&lt;li&gt;modify the structure of auxiliary Hamiltonians,&lt;/li&gt;
&lt;li&gt;combine ideas from different control paradigms,&lt;/li&gt;
&lt;li&gt;generate executable simulation code.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The search space itself becomes dynamic.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Physics solvers provide rigorous validation
&lt;/h3&gt;

&lt;p&gt;Candidate protocols are evaluated through differentiable quantum many-body simulations powered by high-performance quantum software infrastructure such as &lt;strong&gt;TensorCircuit-NG&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The system does not rely on language-model judgment alone—the generated ideas must survive quantitative physical evaluation.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Memory engine turns experiments into reusable knowledge
&lt;/h3&gt;

&lt;p&gt;Every successful or failed experiment is automatically recorded as structured knowledge.&lt;/p&gt;

&lt;p&gt;Over time, the system accumulates reusable design principles, allowing previous discoveries to influence future exploration.&lt;/p&gt;

&lt;p&gt;Through this process, AI evolves from a parameter fitting tool into a continuously improving scientific assistant.&lt;/p&gt;




&lt;h1&gt;
  
  
  Three Levels of Evolution: From Pulse Shaping to Neural Control Generators
&lt;/h1&gt;

&lt;p&gt;To demonstrate the capability of QOC-Workbench, the authors tested it across three increasingly challenging physical scenarios.&lt;/p&gt;

&lt;p&gt;These examples illustrate a gradual transition:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;from optimizing existing protocols → modifying physical pathways → discovering new computational paradigms.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Case 1: Designing Hardware-Compatible Control Pulses
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Rydberg Atom Arrays
&lt;/h3&gt;

&lt;p&gt;The first challenge considers solving the Maximum Independent Set problem using Rydberg atom arrays.&lt;/p&gt;

&lt;p&gt;Real quantum hardware imposes strict constraints on available control signals. Traditional approaches often rely on analytical counterdiabatic protocols derived from simplified models.&lt;/p&gt;

&lt;p&gt;However, these idealized solutions may not fully capture the complexity of interacting many-body systems.&lt;/p&gt;

&lt;p&gt;QOC-Workbench analyzed the limitations of existing approaches and explored a broader control space.&lt;/p&gt;

&lt;p&gt;Instead of simply tuning parameters of known pulses, it discovered a new pulse structure:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;the “smooth beta-bump” envelope.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The generated protocol:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;satisfies realistic hardware constraints,&lt;/li&gt;
&lt;li&gt;preserves smooth control behavior,&lt;/li&gt;
&lt;li&gt;achieves higher ground-state fidelity than classical analytical baselines in many-body simulations.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This demonstrates that LLM-driven exploration can redesign control waveforms rather than merely optimize them.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqyta87ehcrrne5dxnk4g.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqyta87ehcrrne5dxnk4g.png" alt="Agent exploration history" width="799" height="541"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Case 2: Redesigning Many-Body Evolution Paths
&lt;/h2&gt;

&lt;h3&gt;
  
  
  XXZ Spin Chains
&lt;/h3&gt;

&lt;p&gt;The second example moves beyond pulse engineering.&lt;/p&gt;

&lt;p&gt;For XXZ spin chains with complex spectral structures, QOC-Workbench explored the structure of the entire evolution trajectory.&lt;/p&gt;

&lt;p&gt;The system introduced:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a target catalyst Hamiltonian that asymptotically vanishes at the endpoints,&lt;/li&gt;
&lt;li&gt;jointly optimized approximate counterdiabatic corrections,&lt;/li&gt;
&lt;li&gt;nonlinear time schedules combining polynomial and sinusoidal deformations.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The key insight was that optimal control is not only about adjusting individual parameters.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;geometry of the evolution path itself&lt;/strong&gt; can be redesigned.&lt;/p&gt;

&lt;p&gt;By jointly optimizing the Hamiltonian pathway and correction terms, the system discovered improved protocols that would be difficult to obtain through conventional parameter optimization alone.&lt;/p&gt;




&lt;h2&gt;
  
  
  Case 3: Breaking the Scaling Barrier with Neural Generators
&lt;/h2&gt;

&lt;h3&gt;
  
  
  2D Random-Field Ising Models
&lt;/h3&gt;

&lt;p&gt;The third case reveals the most significant conceptual shift.&lt;/p&gt;

&lt;p&gt;For disordered many-body systems, optimizing a control protocol separately for every instance quickly becomes computationally expensive.&lt;/p&gt;

&lt;p&gt;QOC-Workbench identified this bottleneck and changed the problem formulation.&lt;/p&gt;

&lt;p&gt;Instead of searching for an optimal protocol instance by instance, it designed and trained a &lt;strong&gt;graph neural network (GNN) generator&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The generator was trained only on small-scale graph instances but successfully generalized to larger unseen systems.&lt;/p&gt;

&lt;p&gt;It could:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;accurately predict control coefficients,&lt;/li&gt;
&lt;li&gt;generate reasonable evolution paths,&lt;/li&gt;
&lt;li&gt;bypass expensive per-instance variational optimization.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This represents a transition from:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“optimize every problem separately”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;to:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“learn the underlying structure of the solution space.”&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  Conclusion: Let Physicists Return to Physics
&lt;/h1&gt;

&lt;p&gt;QOC-Workbench is not designed to replace human scientific intuition.&lt;/p&gt;

&lt;p&gt;Instead, it aims to amplify it.&lt;/p&gt;

&lt;p&gt;In this emerging human-AI collaboration paradigm, researchers no longer need to spend most of their time on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;manual parameter tuning,&lt;/li&gt;
&lt;li&gt;repetitive protocol benchmarking,&lt;/li&gt;
&lt;li&gt;low-level implementation details.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Instead, they can focus on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;understanding fundamental physical mechanisms,&lt;/li&gt;
&lt;li&gt;defining meaningful physical constraints,&lt;/li&gt;
&lt;li&gt;interpreting machine-discovered protocols,&lt;/li&gt;
&lt;li&gt;extracting new scientific principles.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Human insights then become new knowledge injected back into the system, creating a continuous feedback loop between human reasoning and machine exploration.&lt;/p&gt;

&lt;p&gt;From manually crafting isolated interpolation curves to building a continuously evolving, auditable, and transferable knowledge system, quantum control is moving beyond fixed optimization frameworks.&lt;/p&gt;

&lt;p&gt;The future of quantum control may not be about finding better parameters inside predefined spaces.&lt;/p&gt;

&lt;p&gt;It may be about building intelligent systems capable of discovering entirely new control paradigms.&lt;/p&gt;

&lt;p&gt;Quantum control is entering the era of &lt;strong&gt;autopilot&lt;/strong&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>quantum</category>
    </item>
    <item>
      <title>Training a 1,000-Qubit, 40,000-Parameter Quantum Algorithm with Full Gradients Using TensorCircuit-NG</title>
      <dc:creator>Shixin Zhang</dc:creator>
      <pubDate>Sat, 18 Jul 2026 02:47:36 +0000</pubDate>
      <link>https://dev.to/refractionray/training-a-1000-qubit-40000-parameter-quantum-algorithm-with-full-gradients-using-4pho</link>
      <guid>https://dev.to/refractionray/training-a-1000-qubit-40000-parameter-quantum-algorithm-with-full-gradients-using-4pho</guid>
      <description>&lt;h2&gt;
  
  
  Introduction
&lt;/h2&gt;

&lt;p&gt;If you wanted to simulate 1,000 quantum qubits and compute exact full gradients for 40,000 parameters on a classical computer, how much compute power would you need?&lt;/p&gt;

&lt;p&gt;For a 1,000-qubit, 10-layer circuit, calculating a single state amplitude might have a manageable computational overhead. However, when you attempt to solve the system's Variational Quantum Eigensolver (VQE) and compute all of its parameter gradients, the difficulty skyrockets. At first glance, this sounds like a job that strictly requires a supercomputing cluster.&lt;/p&gt;

&lt;p&gt;The underlying reasons for this explosion in complexity are twofold:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The Introduction of the Hamiltonian&lt;/strong&gt;: Computing the VQE expectation value requires evaluating $E(\boldsymbol\theta)=\langle\psi(\boldsymbol\theta)\vert H\vert\psi(\boldsymbol\theta)\rangle$. This transforms the tensor network from a single-sided state contraction into a "bra-MPO-ket" three-layer sandwich structure, immediately more than doubling the graph size.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Memory Wall in Backpropagation&lt;/strong&gt;: Deriving gradients for tens of thousands of parameters requires backpropagation. To compute these gradients in a contraction graph, the compiler has to keep massive amounts of intermediate results from the forward pass in VRAM. This is vastly more difficult—and puts far more pressure on memory—than a simple forward contraction or amplitude calculation.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;In this technical blog, we'll demonstrate how we cut the per-step execution time of a 1,000-qubit VQE down to just &lt;strong&gt;1 second on a single NVIDIA RTX 6000D GPU&lt;/strong&gt;. We achieve this by exploring three different physics- and engineering-driven contraction strategies, all without sacrificing exact gradient precision.&lt;/p&gt;

&lt;h2&gt;
  
  
  Performance Comparison: Three Physical Perspectives
&lt;/h2&gt;

&lt;p&gt;We use a 1D Transverse Field Ising Model (TFIM) with open boundary conditions as our benchmark. Each layer of the circuit applies diagonal entangling gates sequentially to adjacent pairs $(0,1),(1,2),\ldots,(n-2,n-1)$ along the qubit chain:&lt;/p&gt;

&lt;p&gt;$$R_{ZZ}(\theta)=\exp(-i\theta Z\otimes Z/2)$$&lt;/p&gt;

&lt;p&gt;Followed by single-qubit rotations. The TFIM Hamiltonian is:&lt;/p&gt;

&lt;p&gt;$$H_{\rm TFIM}=\sum_{i=0}^{n-2}X_iX_{i+1}+\sum_{i=0}^{n-1}Z_i$$&lt;/p&gt;

&lt;p&gt;We can also represent this Hamiltonian as an exact and compact Matrix Product Operator (MPO, with a Bond Dimension of 3).&lt;/p&gt;

&lt;p&gt;The table below shows the benchmark results for three different contraction modes (using an NVIDIA RTX 6000D at &lt;code&gt;complex64&lt;/code&gt; precision):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;strong&gt;Perspective (Algorithm Mode)&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Initial Compilation&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;First Full Gradient&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Steady-State Full Gradient&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Peak VRAM&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Applicability &amp;amp; Value&lt;/strong&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Global Contraction Tree&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;5 h 20 min&lt;/td&gt;
&lt;td&gt;5.88 s&lt;/td&gt;
&lt;td&gt;1.09 s&lt;/td&gt;
&lt;td&gt;51.97 GB&lt;/td&gt;
&lt;td&gt;Directly contracts the full bra-MPO-ket graph; useful for full-graph compute and cross-benchmarking.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Local Causal Sliding Window&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;10.79 s&lt;/td&gt;
&lt;td&gt;148.68 s&lt;/td&gt;
&lt;td&gt;148.36 s&lt;/td&gt;
&lt;td&gt;8.86 GB&lt;/td&gt;
&lt;td&gt;Relies on strict causal cones of fixed width for local terms; terms can be dispatched independently.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Local Window (8-GPU Parallel)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;10.50–11.28 s / GPU&lt;/td&gt;
&lt;td&gt;19.72–20.50 s&lt;/td&gt;
&lt;td&gt;18.71 s&lt;/td&gt;
&lt;td&gt;8.86 GB / GPU&lt;/td&gt;
&lt;td&gt;Dispatches 125 local terms per GPU; yielded a measured 7.93× speedup.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Spatial Transfer Matrix (B3)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;6.367 s&lt;/td&gt;
&lt;td&gt;1.328 s&lt;/td&gt;
&lt;td&gt;1.054 s&lt;/td&gt;
&lt;td&gt;8.62 GB&lt;/td&gt;
&lt;td&gt;Leverages repeating 1D spatial structures.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Spatial Transfer Matrix (B5)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;7.375 s&lt;/td&gt;
&lt;td&gt;1.320 s&lt;/td&gt;
&lt;td&gt;1.133 s&lt;/td&gt;
&lt;td&gt;5.44 GB&lt;/td&gt;
&lt;td&gt;A lower-VRAM variant of the same exact STM algorithm.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These approaches complement one another:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The Global Tree&lt;/strong&gt; delivers excellent steady-state execution time (1.09s) but suffers from a brutal 5+ hour cold compilation time and massive 52 GB peak VRAM usage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Local Window&lt;/strong&gt; approach crushes compile time down to ~10 seconds, but redundantly computes the shared environments of adjacent local terms, making single-GPU execution sluggish.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Spatial Transfer Matrix (STM)&lt;/strong&gt; successfully decouples compilation costs from the system length by explicitly passing the repetitive 1D spatial structure to &lt;code&gt;jax.lax.scan&lt;/code&gt;, perfectly balancing compilation and execution efficiency.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa4ko4snbtifzomsp4l11.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa4ko4snbtifzomsp4l11.png" alt="scatter plot for efficiency of three methods" width="800" height="591"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Contraction Path Metrics: FLOPs, Max Tensor, and Write
&lt;/h2&gt;

&lt;p&gt;In tensor network contraction, path quality is defined by several resource metrics:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;FLOPs (Floating Point Operations)&lt;/strong&gt;: The total number of scalar multiply-add operations required for the path. For example, &lt;code&gt;log10 FLOPs = 12.0&lt;/code&gt; means roughly $10^{12}$ scalar ops.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Max Tensor Size (Width $w$)&lt;/strong&gt;: The number of elements in the largest intermediate tensor generated during contraction. It directly dictates the instantaneous VRAM pressure during the forward pass. (If all indices have a dimension of 2, width $w$ corresponds to $2^w$ elements).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Total Write&lt;/strong&gt;: The total number of elements written out across all intermediate tensors. During backpropagation, &lt;em&gt;all&lt;/em&gt; of these written tensors must be retained in memory as residuals to compute gradients.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When calculating expectations and gradients simultaneously, you must keep an eye on all three metrics. Ultimately, however, empirical compile time, execution time, and peak GPU VRAM are the ground truth.&lt;/p&gt;

&lt;h2&gt;
  
  
  Perspective 1: Global Contraction Tree — Brute-Forcing the Entire Graph
&lt;/h2&gt;

&lt;p&gt;By calling &lt;code&gt;value_and_grad&lt;/code&gt; directly on the complete bra-MPO-ket network, we can quantify the true cost of naive global contraction.&lt;/p&gt;

&lt;p&gt;To make this feasible, we first alter the network representation. It is critical to write each $R_{ZZ}$ gate in an exact Rank-2 decomposed form:&lt;/p&gt;

&lt;p&gt;$$R_{ZZ}(\theta)=\cos(\theta/2)I\otimes I-i\sin(\theta/2)Z\otimes Z$$&lt;/p&gt;

&lt;p&gt;Maintaining the original left-to-right "ladder" gate ordering is also crucial. If we reorder the commuting $R_{ZZ}$ gates into an alternating even/odd bond "brick-wall" structure, we destroy the elimination path. The tensor network width instantly spikes to 39.585. Keeping the ladder ordering keeps the width at a manageable 23.585.&lt;/p&gt;

&lt;p&gt;Using the &lt;code&gt;omeco&lt;/code&gt; path search on this massive graph, we found a global contraction path with &lt;code&gt;log10 FLOPs = 12.0466&lt;/code&gt; and &lt;code&gt;log2 write = 32.5912&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;On a single RTX 6000D, post-compilation gradient execution takes only 1.09s, peaking at 51 GB of VRAM. This proves that full-graph VQE gradients can technically be run on a single card, but the bottlenecks are purely the global HLO compilation time and the memory needed to store backward residuals.&lt;/p&gt;

&lt;h2&gt;
  
  
  omeco vs. cotengra: Two Path-Search Frameworks
&lt;/h2&gt;

&lt;p&gt;Since the global graph's resource bottleneck lies heavily in contraction path search and compilation, how do top-tier tools fare? We compared two frameworks—&lt;code&gt;omeco&lt;/code&gt; and &lt;code&gt;cotengra&lt;/code&gt;—on this massive tensor network.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;omeco&lt;/strong&gt; (Rust-based) uses a TreeSA algorithm to perform intense simulated annealing in tree space. It is excellent for quickly finding high-quality un-sliced seed trees on large graphs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;cotengra&lt;/strong&gt; offers richer interfaces for manipulating and fine-tuning contraction trees, such as reconfiguring subtrees based on existing paths or performing granular slice-index searches on a fixed tree.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We ran a controlled search comparison on a 100-qubit, 10-layer TFIM graph (6,460 tensors, 8,441 indices):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;strong&gt;Search Method &amp;amp; Budget&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Search Time&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;log10 FLOPs&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;log2 Max Tensor&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;log2 Total Write&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Notes&lt;/strong&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;omeco&lt;/code&gt; TreeSA (16 trials × 64 steps)&lt;/td&gt;
&lt;td&gt;15.68 s&lt;/td&gt;
&lt;td&gt;11.068&lt;/td&gt;
&lt;td&gt;24.585&lt;/td&gt;
&lt;td&gt;29.517&lt;/td&gt;
&lt;td&gt;16 independent SA trials&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;cotengra&lt;/code&gt; CMA-ES (8 hyperparam candidates)&lt;/td&gt;
&lt;td&gt;37.59 s&lt;/td&gt;
&lt;td&gt;12.901&lt;/td&gt;
&lt;td&gt;31.585&lt;/td&gt;
&lt;td&gt;35.646&lt;/td&gt;
&lt;td&gt;8 sets of search params&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;cotengra&lt;/code&gt; SA (12 steps × 12 reconfigs)&lt;/td&gt;
&lt;td&gt;9.22 s&lt;/td&gt;
&lt;td&gt;11.827&lt;/td&gt;
&lt;td&gt;29.585&lt;/td&gt;
&lt;td&gt;34.014&lt;/td&gt;
&lt;td&gt;Low-budget SA&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;cotengra&lt;/code&gt; SA (24 steps × 24 reconfigs)&lt;/td&gt;
&lt;td&gt;34.09 s&lt;/td&gt;
&lt;td&gt;10.873&lt;/td&gt;
&lt;td&gt;23.585&lt;/td&gt;
&lt;td&gt;30.367&lt;/td&gt;
&lt;td&gt;Higher-budget SA&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The data shows that &lt;code&gt;omeco&lt;/code&gt; TreeSA ran significantly more annealing steps in less time and returned lower-complexity trees, outperforming the alternatives for this un-sliced scenario. However, for scenarios requiring deep slicing, neither framework is currently ideal.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Efficiency Limits of Slicing
&lt;/h2&gt;

&lt;p&gt;When VRAM is tight, slicing the contraction tree is a necessary evil to keep memory footprints in check. But if you have enough VRAM for an un-sliced path, forcing a slice to distribute across multiple GPUs does not scale efficiency linearly.&lt;/p&gt;

&lt;p&gt;Fundamentally, slicing a tensor network means forcibly delaying the contraction of selected indices to the very end of the path. In an optimal, un-sliced path, these indices are eliminated early on, preventing intermediate tensors from blowing up in size. Slicing disrupts this optimal elimination order. It forces every sub-task to redundantly compute intermediate results that could have been shared before the final summation.&lt;/p&gt;

&lt;p&gt;When we force-sliced the 1,000-qubit global path into 8, 32, and 128 tasks, the &lt;code&gt;log2 write&lt;/code&gt; per slice barely dropped—from 32.5912 to 32.5632, 32.5491, and 32.5365 respectively. Because it failed to meaningfully dismantle the primary memory structure of the backward pass, while heavily deviating from the optimal contraction order, redundant computation skyrocketed. Total &lt;code&gt;log10 FLOPs&lt;/code&gt; jumped from 12.0466 to 12.9420, 13.5391, and 14.1355.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Takeaway:&lt;/strong&gt; If it fits in VRAM, using an un-sliced path usually maximizes hardware utilization.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5o3kezbmdg3v1djilwap.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5o3kezbmdg3v1djilwap.png" alt="framework of the three methods" width="799" height="270"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Perspective 2: Local Causal Sliding Window — Exploiting Local Physics
&lt;/h2&gt;

&lt;p&gt;For a circuit depth of $L=10$, we can leverage the commutativity of 2-qubit gates to convert the ladder structure into a brick-wall equivalent. Consequently, every local TFIM term possesses an exact backward causal cone during backpropagation. Since the diagonal entangling gates commute, the width of this causal window is strictly $2L+2$, which is only 22 sites for $L=10$. The brick-wall structure, which was a liability for global tree search, becomes a massive advantage here.&lt;/p&gt;

&lt;p&gt;By using &lt;code&gt;lax.scan&lt;/code&gt; to iterate over the local Pauli terms of the Hamiltonian and wrapping single terms in &lt;code&gt;jax.checkpoint&lt;/code&gt;, the compilation scale becomes solely dependent on window size and circuit depth, totally decoupled from the total number of system sites:&lt;/p&gt;

&lt;p&gt;$$T_{\rm window}=O(n C_{\rm local}(L)),\qquad M_{\rm window}=O(M_{\rm local}(L))$$&lt;/p&gt;

&lt;p&gt;This approach boasts rapid compile times (~10 seconds) and a lean VRAM footprint of 8.86 GB. If dispatched across 8 GPUs (125 local terms per GPU), the steady-state gradient time drops to 18.71s. However, the overlapping causal windows introduce massive computational redundancy.&lt;/p&gt;

&lt;p&gt;Furthermore, this method heavily relies on the commutativity of diagonal gates. If we swap them for non-commuting 2-qubit gates, the causal cone under a ladder structure expands to $O(n)$, rendering this method useless.&lt;/p&gt;

&lt;h2&gt;
  
  
  Perspective 3: Spatial Transfer Matrix — Lossless Exact Scanning
&lt;/h2&gt;

&lt;p&gt;The spatial transfer matrix strategy slices the system along the spatial qubit dimension into a left-end block, repeating middle blocks, and a right-end block. The middle block receives the complete boundary tensor from the left, contracts its internal bra-MPO-ket subnetwork, and outputs to the next block.&lt;/p&gt;

&lt;p&gt;The recurrence relation is:&lt;/p&gt;

&lt;p&gt;$$B_{k+1}=\mathcal{T}_k(\boldsymbol\theta_k)B_k$$&lt;/p&gt;

&lt;p&gt;At depth $L=10$, the spatial cut crosses 10 ket bonds, 10 bra bonds, and 1 MPO bond (dimension 3). The exact boundary size is:&lt;/p&gt;

&lt;p&gt;$$D_{\rm boundary}=3\times2^{2L}=3\times4^L = 3\times4^{10}$$&lt;/p&gt;

&lt;p&gt;This amounts to roughly 24 MiB of memory traffic, which perfectly explains why this method can effortlessly perform linear scanning at the 1,000-qubit scale.&lt;/p&gt;

&lt;p&gt;Using a 3-site block division, this mode requires just 6.367s for cold compilation, executes at 1.054s per step, and peaks at 8.62 GB of VRAM.&lt;/p&gt;

&lt;p&gt;The spatial width of each scan step (the block size $b$) acts as a tunable time-space tradeoff:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Smaller blocks&lt;/strong&gt; shrink the internal contraction graph, reducing compile pressure and the residuals saved for backprop, but increase the number of scan steps needed to pass the boundary tensor.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Larger blocks&lt;/strong&gt; reduce scan steps but inflate maximum tensor size, total write, compile time, and peak VRAM.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In our tests, a 3-site block offered optimal throughput. A 5-site block reduced peak VRAM further to 5.44 GB, at the slight cost of increasing execution time to 1.133s.&lt;/p&gt;

&lt;h2&gt;
  
  
  Counterintuitive Acceleration: Trading Compute for Bandwidth
&lt;/h2&gt;

&lt;p&gt;Fascinatingly, pairing the transfer matrix's middle block loop with &lt;code&gt;jax.checkpoint&lt;/code&gt; (recomputation) dropped the peak VRAM from 42.24 GB down to 8.62 GB, while &lt;em&gt;also&lt;/em&gt; slightly reducing steady-state execution time from 1.118s to 1.054s.&lt;/p&gt;

&lt;p&gt;"Reducing memory traffic by recomputing intermediate residuals" actually speeding up execution highlights a hardware reality: on modern consumer GPUs, memory bandwidth is often a much harder bottleneck than raw single-precision FLOPs. Taking on a heavier compute burden to alleviate read/write strain yields net performance gains. This same logic applies when tuning the penalty ratio between time and space complexity during contraction tree searches.&lt;/p&gt;

&lt;p&gt;Even more thought-provoking is the fact that this relatively simple spatial slicing and block-scanning strategy heavily outperformed the global tree—derived from intense simulated annealing—in execution time, compile time, and memory footprint. This implies that for truly massive tensor networks requiring backpropagation (where intermediate residuals and VRAM are paramount), existing contraction path algorithms still have significant blind spots and ample room for innovation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pitfalls to Avoid
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Bypassing TF32 Precision Limitations&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;When doing complex matrix multiplications on NVIDIA GPUs, simply setting the highest precision in JAX is not enough to avoid automatic degradation to TF32. You must explicitly set this environment variable before runtime:&lt;/p&gt;

&lt;p&gt;Bash&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;   export NVIDIA_TF32_OVERRIDE=0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you don't, the state norm can slip to 0.98–0.996, causing noticeable deviations in energy calculations. This is a stark reminder of the differing precision tolerances between standard Machine Learning and Quantum Simulation.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Rayon Thread Pool Stack Overflows&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;When &lt;code&gt;omeco&lt;/code&gt; searches ultra-deep trees (depth &amp;gt; 1000) and generates slicing code, the recursive traversal of the deep tree can trigger a Segmentation Fault in Rayon threads. You must explicitly bump up the stack size before initiating Python or the Rayon global pool:&lt;/p&gt;

&lt;p&gt;Bash&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;   export RUST_MIN_STACK=67108864
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Full 1,000-Qubit VQE: From Single-Step Gradients to End-to-End Optimization
&lt;/h2&gt;

&lt;p&gt;To prove that this isn't just a parlor trick for a single energy/gradient calculation, we ran a complete, complex 1,000-qubit VQE optimization. Our circuit features 10 layers of nearest-neighbor $R_{ZZ}$ ladders, with each layer's single-qubit component consisting of sequential $R_X$, $R_Z$, and $R_X$ rotations:&lt;/p&gt;

&lt;p&gt;$$\vert{}0\rangle^{\otimes 1000}\xrightarrow{H^{\otimes 1000}}\left[R_{ZZ}\text{ Ladder}\rightarrow R_X\rightarrow R_Z\rightarrow R_X\right]^{10}.$$&lt;/p&gt;

&lt;p&gt;Before constructing the full bra-MPO-ket network, we exactly compressed each consecutive $R_X$–$R_Z$–$R_X$ sequence into a single $2\times2$ single-qubit tensor. We then utilized the exact same Spatial Transfer Matrix scheme (with zero approximation truncations) to compute the energy and the gradients for all 40,000 parameters. The entire optimization ran on a single NVIDIA RTX 6000D, completing 2,000 Adam updates.&lt;/p&gt;

&lt;p&gt;For this 1,000-site critical open-boundary TFIM, the exact ground state energy can be derived directly from the free-fermion spectrum. After 2,000 updates, our VQE energy was $E_{\rm VQE}=-1268.4365$, resulting in a total energy error of $E_{\rm VQE}-E_0=4.440$. This equates to a relative error of about $3\times10^{-3}$.&lt;/p&gt;

&lt;p&gt;Without any hyperparameter tuning, a 1,000-qubit VQE featuring complex gate sequences, full parameters, and exact full gradients hit a relative energy precision in the thousandths. This confirms that our tensor network representations, path searches, spatial scanning, autodiff, and GPU execution aren't isolated micro-optimizations—they form a cohesive, end-to-end technical pipeline capable of translating circuit definitions into actual, strictly-verified optimization convergence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Our work on full-gradient simulation and optimization for a 1,000-qubit VQE demonstrates that by hybridizing low-rank gate decomposition, contraction path search, lossless spatial transfer matrices, JAX automatic differentiation, and compiler mechanics, we can do the impossible. We can compute the full gradient of a 1,000-qubit VQE rapidly on a single GPU, execute thousands of training steps within hours, and achieve ~0.3% relative energy error without systemic fine-tuning.&lt;/p&gt;

&lt;p&gt;These results illustrate the immense potential of moving from single-shot simulations to massive end-to-end variational computation, leaving plenty of headroom for exploring better circuit structures, learning rate schedules, and higher precision.&lt;/p&gt;

&lt;p&gt;Getting a 1,000-qubit VQE to run natively proves one thing: in the deep waters of quantum simulation, the mathematical elegance of algorithmic design must be deeply coupled with the underlying mechanics of compilers (XLA), automatic differentiation, and GPU hardware architectures. TensorCircuit-NG is more than just a simulator; it is a critical piece of infrastructure that bridges software and hardware, providing a complete technical pipeline from quantum algorithm definition to ultra-large-scale variational optimization.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>jax</category>
      <category>quantum</category>
    </item>
    <item>
      <title>Large-Scale TensorCircuit Contractions on GPUs: Disabling XLA GPU Autotuning</title>
      <dc:creator>Shixin Zhang</dc:creator>
      <pubDate>Wed, 15 Jul 2026 03:41:23 +0000</pubDate>
      <link>https://dev.to/refractionray/large-scale-tensorcircuit-contractions-on-gpus-disabling-xla-gpu-autotuning-3p2</link>
      <guid>https://dev.to/refractionray/large-scale-tensorcircuit-contractions-on-gpus-disabling-xla-gpu-autotuning-3p2</guid>
      <description>&lt;p&gt;When running large-scale tensor-network contractions with TensorCircuit-NG and the JAX GPU backend, the following runtime configuration is worth testing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;XLA_PYTHON_CLIENT_PREALLOCATE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;false &lt;/span&gt;&lt;span class="nv"&gt;XLA_FLAGS&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nt"&gt;--xla_gpu_autotune_level&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;0 python your_script.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Its main benefit is not speed, but lower persistent GPU memory usage from XLA GPU autotuning, which makes memory behavior during compilation and on the first visible GPU more predictable. In the TensorCircuit contraction workloads we tested, disabling autotuning also slightly improved steady-state runtime, but the memory savings were the more important result.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key takeaway
&lt;/h2&gt;

&lt;p&gt;XLA GPU autotuning evaluates alternative algorithms or workspace configurations for certain GPU kernels and custom calls, then selects an implementation. This can be valuable for convolutions, large GEMMs, and deep-learning workloads with fixed shapes. For large TensorCircuit contractions, however, the contraction path is already determined by OMECO or cotengra, leaving relatively little optimization freedom for autotuning while still potentially incurring substantial persistent memory overhead during compilation and tuning.&lt;/p&gt;

&lt;p&gt;For TensorCircuit contractions, run this A/B test by default:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Baseline&lt;/span&gt;
&lt;span class="nv"&gt;XLA_PYTHON_CLIENT_PREALLOCATE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;false &lt;/span&gt;python your_script.py

&lt;span class="c"&gt;# Test configuration&lt;/span&gt;
&lt;span class="nv"&gt;XLA_PYTHON_CLIENT_PREALLOCATE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;false &lt;/span&gt;&lt;span class="nv"&gt;XLA_FLAGS&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nt"&gt;--xla_gpu_autotune_level&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;0 python your_script.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both environment variables must be set before the Python process starts and before JAX is imported. This recommendation primarily concerns GPUs; CPU backends do not exhibit the same GPU-kernel autotuning behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  Representative results
&lt;/h2&gt;

&lt;p&gt;All results below use a fixed contraction path so that path-search randomness does not affect the comparison.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Workload&lt;/th&gt;
&lt;th&gt;Autotuning&lt;/th&gt;
&lt;th&gt;Post-compile memory&lt;/th&gt;
&lt;th&gt;Peak memory&lt;/th&gt;
&lt;th&gt;Steady-state runtime&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;100 qubits × 24 layers, amplitude&lt;/td&gt;
&lt;td&gt;Default&lt;/td&gt;
&lt;td&gt;8.7 GiB&lt;/td&gt;
&lt;td&gt;8.7 GiB&lt;/td&gt;
&lt;td&gt;0.37 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;100 qubits × 24 layers, amplitude&lt;/td&gt;
&lt;td&gt;autotune=0&lt;/td&gt;
&lt;td&gt;0.5 GiB&lt;/td&gt;
&lt;td&gt;4.6 GiB&lt;/td&gt;
&lt;td&gt;0.40 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;28 qubits × 12 layers, expectation gradient&lt;/td&gt;
&lt;td&gt;Default&lt;/td&gt;
&lt;td&gt;4.7 GiB&lt;/td&gt;
&lt;td&gt;21.4 GiB&lt;/td&gt;
&lt;td&gt;0.18 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;28 qubits × 12 layers, expectation gradient&lt;/td&gt;
&lt;td&gt;autotune=0&lt;/td&gt;
&lt;td&gt;0.5 GiB&lt;/td&gt;
&lt;td&gt;17.2 GiB&lt;/td&gt;
&lt;td&gt;0.15 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;28 qubits × 13 layers, expectation gradient&lt;/td&gt;
&lt;td&gt;Default&lt;/td&gt;
&lt;td&gt;10.9 GiB&lt;/td&gt;
&lt;td&gt;43.8 GiB&lt;/td&gt;
&lt;td&gt;0.46 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;28 qubits × 13 layers, expectation gradient&lt;/td&gt;
&lt;td&gt;autotune=0&lt;/td&gt;
&lt;td&gt;0.5 GiB&lt;/td&gt;
&lt;td&gt;33.5 GiB&lt;/td&gt;
&lt;td&gt;0.42 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;28 qubits × 14 layers, expectation gradient&lt;/td&gt;
&lt;td&gt;Default&lt;/td&gt;
&lt;td&gt;17.0 GiB&lt;/td&gt;
&lt;td&gt;64.5 GiB&lt;/td&gt;
&lt;td&gt;0.58 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;28 qubits × 14 layers, expectation gradient&lt;/td&gt;
&lt;td&gt;autotune=0&lt;/td&gt;
&lt;td&gt;0.5 GiB&lt;/td&gt;
&lt;td&gt;64.5 GiB&lt;/td&gt;
&lt;td&gt;0.58 s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The first three workload groups show the most common benefit: disabling autotuning reduces compilation-stage memory and the final peak. The 28 × 14 case appears different: post-compile memory falls from 17.0 GiB to 0.5 GiB, but the final peak is unchanged. This behavior is related to the JAX GPU allocator.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the final peak is unchanged for the 28 × 14 case
&lt;/h2&gt;

&lt;p&gt;Even with &lt;code&gt;XLA_PYTHON_CLIENT_PREALLOCATE=false&lt;/code&gt;, the default JAX GPU allocator tends to retain and reuse GPU memory that it has already allocated. The memory reported by &lt;code&gt;nvidia-smi&lt;/code&gt; is therefore the amount held by the process, not the total size of tensors that are currently live.&lt;/p&gt;

&lt;p&gt;In the 28 × 14 example, default autotuning did consume approximately 17 GiB of additional memory during compilation. During the first real backward contraction, however, the runtime buffers themselves also required a large allocation. The default allocator could reuse blocks allocated earlier, so the final peak was not simply the sum of runtime memory and autotuning memory.&lt;/p&gt;

&lt;p&gt;Using &lt;code&gt;XLA_PYTHON_CLIENT_ALLOCATOR=platform&lt;/code&gt; as a diagnostic exposes this difference:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Workload&lt;/th&gt;
&lt;th&gt;Allocator&lt;/th&gt;
&lt;th&gt;Autotuning&lt;/th&gt;
&lt;th&gt;Post-compile memory&lt;/th&gt;
&lt;th&gt;Peak memory&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;28 qubits × 14 layers, expectation gradient&lt;/td&gt;
&lt;td&gt;Default&lt;/td&gt;
&lt;td&gt;Default&lt;/td&gt;
&lt;td&gt;17.0 GiB&lt;/td&gt;
&lt;td&gt;64.5 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;28 qubits × 14 layers, expectation gradient&lt;/td&gt;
&lt;td&gt;Default&lt;/td&gt;
&lt;td&gt;autotune=0&lt;/td&gt;
&lt;td&gt;0.5 GiB&lt;/td&gt;
&lt;td&gt;64.5 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;28 qubits × 14 layers, expectation gradient&lt;/td&gt;
&lt;td&gt;platform&lt;/td&gt;
&lt;td&gt;Default&lt;/td&gt;
&lt;td&gt;17.0 GiB&lt;/td&gt;
&lt;td&gt;82.8 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;28 qubits × 14 layers, expectation gradient&lt;/td&gt;
&lt;td&gt;platform&lt;/td&gt;
&lt;td&gt;autotune=0&lt;/td&gt;
&lt;td&gt;0.5 GiB&lt;/td&gt;
&lt;td&gt;66.3 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The &lt;code&gt;platform&lt;/code&gt; allocator retains less reusable GPU memory, which makes it useful for diagnosing where allocations originate. Because it also reuses fewer large blocks allocated earlier, however, the runtime may request additional memory and produce a higher peak. Prefer the default allocator for normal execution; use &lt;code&gt;platform&lt;/code&gt; mainly to confirm whether autotuning introduces extra memory usage.&lt;/p&gt;

&lt;h2&gt;
  
  
  Interpreting peak memory
&lt;/h2&gt;

&lt;p&gt;Here, peak memory is the maximum per-process GPU memory observed by polling &lt;code&gt;nvidia-smi&lt;/code&gt;, with snapshots recorded at stages such as &lt;code&gt;after_compile&lt;/code&gt; and &lt;code&gt;after_first_run&lt;/code&gt;. XLA's &lt;code&gt;compiled.memory_analysis().peak_memory_in_bytes&lt;/code&gt; measures something different: it more closely reflects the computation graph's buffer assignment and is usually lower than the actual process memory reported by &lt;code&gt;nvidia-smi&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the first GPU may use more memory in multi-GPU processes
&lt;/h2&gt;

&lt;p&gt;In a JAX/XLA process that can see multiple GPUs, backend initialization, the default device, compilation services, autotuning caches, executable caches, or allocator state may be placed preferentially on the first GPU in the visible-device list. Consequently, the first GPU can hold an extra block of memory. Disabling &lt;code&gt;xla_gpu_autotune_level&lt;/code&gt; often reduces this first-device memory tax.&lt;/p&gt;

&lt;p&gt;For multiple independent single-GPU jobs, use &lt;code&gt;CUDA_VISIBLE_DEVICES&lt;/code&gt; so that each process sees only its assigned GPU:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;CUDA_VISIBLE_DEVICES&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1 &lt;span class="nv"&gt;XLA_PYTHON_CLIENT_PREALLOCATE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;false &lt;/span&gt;&lt;span class="nv"&gt;XLA_FLAGS&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nt"&gt;--xla_gpu_autotune_level&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;0 python your_script.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For genuine multi-GPU parallel or distributed jobs, do not hide GPUs that must participate in the computation merely to avoid the first-device memory tax. Preserve the correct device set, then measure how disabling autotuning changes memory usage on each GPU.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practical recommendations
&lt;/h2&gt;

&lt;p&gt;For large TensorCircuit-NG contractions, use the following procedure:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Fix the contraction path instead of running a new stochastic path search for every benchmark.&lt;/li&gt;
&lt;li&gt;Set &lt;code&gt;XLA_PYTHON_CLIENT_PREALLOCATE=false&lt;/code&gt; to prevent JAX from reserving a large block of GPU memory at startup.&lt;/li&gt;
&lt;li&gt;A/B test default autotuning against &lt;code&gt;XLA_FLAGS=--xla_gpu_autotune_level=0&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Record compile time, first-run time, steady-state runtime, post-compile memory, and post-first-run memory separately.&lt;/li&gt;
&lt;li&gt;Even if disabling autotuning does not improve steady-state runtime, prefer &lt;code&gt;autotune=0&lt;/code&gt;, especially near the OOM limit or when scheduling multiple GPUs.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;In short, disabling XLA GPU autotuning does not change the contraction path; it reduces an additional source of memory usage in the GPU compilation and execution layers. For large TensorCircuit contractions, this is usually a low-risk configuration well worth testing.&lt;/p&gt;

</description>
      <category>jax</category>
    </item>
    <item>
      <title>When Quantum Dynamics Doesn't Start from Zero: Completing the Missing Half of Entanglement Growth | PRL</title>
      <dc:creator>Shixin Zhang</dc:creator>
      <pubDate>Fri, 10 Jul 2026 00:42:59 +0000</pubDate>
      <link>https://dev.to/refractionray/when-quantum-dynamics-doesnt-start-from-zero-completing-the-missing-half-of-entanglement-growth--pl7</link>
      <guid>https://dev.to/refractionray/when-quantum-dynamics-doesnt-start-from-zero-completing-the-missing-half-of-entanglement-growth--pl7</guid>
      <description>&lt;p&gt;Entanglement dynamics lies at the heart of nonequilibrium quantum physics. For more than two decades, the standard approach has been remarkably consistent: start from an unentangled product state, let the system evolve unitarily, and study how the entanglement entropy grows over time.&lt;/p&gt;

&lt;p&gt;This framework has led to many fundamental discoveries, including our understanding of quantum thermalization and many-body localization (MBL). It also shaped an implicit assumption shared across the field: whenever the half-chain entanglement entropy increases, the system must be &lt;em&gt;creating&lt;/em&gt; new quantum entanglement.&lt;/p&gt;

&lt;p&gt;Our recent paper, published in &lt;em&gt;Physical Review Letters&lt;/em&gt;, argues that this picture is incomplete.&lt;/p&gt;

&lt;p&gt;The central observation is surprisingly simple:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;An increase in measured entanglement does not necessarily mean new entanglement has been created.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Instead, quantum dynamics possesses two fundamentally different capabilities that have long been mixed together.&lt;/p&gt;




&lt;h2&gt;
  
  
  Two ways to increase entanglement
&lt;/h2&gt;

&lt;p&gt;A useful way to think about entanglement is to imagine the system as a connected network of water reservoirs.&lt;/p&gt;

&lt;p&gt;There are two distinct mechanisms that can raise the water level observed at a particular cut.&lt;/p&gt;

&lt;h3&gt;
  
  
  Build: Creating new entanglement
&lt;/h3&gt;

&lt;p&gt;The first mechanism genuinely generates new quantum correlations.&lt;/p&gt;

&lt;p&gt;This is analogous to pumping fresh water into the entire reservoir system. The total amount of entanglement increases.&lt;/p&gt;

&lt;h3&gt;
  
  
  Move: Transporting existing entanglement
&lt;/h3&gt;

&lt;p&gt;The second mechanism creates &lt;strong&gt;no new entanglement at all&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Instead, quantum evolution simply redistributes the entanglement that already exists inside the system. Water is flowing internally, but no new water enters the reservoirs. Although the total amount remains unchanged, the water level measured at one particular location may still rise.&lt;/p&gt;

&lt;p&gt;Exactly the same phenomenon can happen for half-chain entanglement entropy.&lt;/p&gt;

&lt;p&gt;This distinction turns out to be much more important than previously appreciated.&lt;/p&gt;




&lt;h2&gt;
  
  
  The intuitive expectation
&lt;/h2&gt;

&lt;p&gt;Once these two mechanisms are separated, an intuitive principle emerges.&lt;/p&gt;

&lt;p&gt;If a system already starts with a large amount of entanglement, there is less room left to generate additional entanglement later.&lt;/p&gt;

&lt;p&gt;Indeed, this is exactly what happens in familiar systems such as&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;quantum chaotic (thermalizing) systems,&lt;/li&gt;
&lt;li&gt;free-fermion systems,&lt;/li&gt;
&lt;li&gt;random quantum circuits with strong scrambling.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;As the initial entanglement increases, the additional entanglement generated during evolution decreases monotonically.&lt;/p&gt;

&lt;p&gt;This is precisely what one would expect if entanglement growth were primarily driven by &lt;strong&gt;building&lt;/strong&gt; new entanglement.&lt;/p&gt;




&lt;h2&gt;
  
  
  MBL breaks the rule
&lt;/h2&gt;

&lt;p&gt;Many-body localized systems tell a completely different story.&lt;/p&gt;

&lt;p&gt;Instead of decreasing monotonically, the entanglement growth exhibits a striking &lt;strong&gt;bell-shaped dependence&lt;/strong&gt; on the initial entanglement.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Starting from nearly product states, the growth is very small.&lt;/li&gt;
&lt;li&gt;Starting from nearly maximally entangled states, the growth is again very small.&lt;/li&gt;
&lt;li&gt;The largest increase occurs at &lt;strong&gt;intermediate initial entanglement&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This behavior is difficult to explain if entanglement growth comes solely from generating new entanglement.&lt;/p&gt;

&lt;p&gt;Something else must be happening.&lt;/p&gt;




&lt;h2&gt;
  
  
  A "pure transport" experiment
&lt;/h2&gt;

&lt;p&gt;To isolate the missing ingredient, we designed a particularly simple reference model.&lt;/p&gt;

&lt;p&gt;Instead of using an interacting Hamiltonian, we considered a random circuit composed &lt;strong&gt;only of SWAP gates&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;This circuit has a remarkable property:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;it &lt;strong&gt;cannot create any entanglement&lt;/strong&gt;, and&lt;/li&gt;
&lt;li&gt;it only exchanges the locations of quantum states.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In other words, it performs &lt;strong&gt;pure entanglement transport&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Surprisingly, the entanglement-growth curve produced by this SWAP-only circuit closely matches the behavior observed in MBL systems, both qualitatively and even quantitatively.&lt;/p&gt;

&lt;p&gt;This comparison reveals the underlying physics.&lt;/p&gt;

&lt;p&gt;MBL is not particularly good at generating new entanglement.&lt;/p&gt;

&lt;p&gt;Instead, it excels at &lt;strong&gt;moving around the entanglement that already exists&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;When the system begins with almost no entanglement, there is simply nothing to transport.&lt;/p&gt;

&lt;p&gt;When it begins nearly saturated, there is no remaining room for rearrangement.&lt;/p&gt;

&lt;p&gt;Only at intermediate entanglement do both ingredients coexist, allowing transport to produce the largest observable increase.&lt;/p&gt;




&lt;h2&gt;
  
  
  Looking beyond a single bipartition
&lt;/h2&gt;

&lt;p&gt;Most previous studies monitor only one quantity: the entanglement across the middle cut of the system.&lt;/p&gt;

&lt;p&gt;While extremely useful, this provides only a partial view of the system's entanglement structure.&lt;/p&gt;

&lt;p&gt;In our work, we introduce a more global perspective by considering the &lt;strong&gt;Bipartition-Averaged Entanglement Entropy (BAEE)&lt;/strong&gt;, which averages entanglement over all possible bipartitions.&lt;/p&gt;

&lt;p&gt;The simulations reveal an interesting phenomenon.&lt;/p&gt;

&lt;p&gt;Even in ordinary thermalizing systems, BAEE grows much faster than the half-chain entanglement during the early stages of evolution.&lt;/p&gt;

&lt;p&gt;The difference between these two quantities represents a hidden reservoir of entanglement that has already been generated locally but has not yet reached the particular cut being measured.&lt;/p&gt;

&lt;p&gt;Returning to our water analogy, thermalization rapidly fills many local reservoirs throughout the system.&lt;/p&gt;

&lt;p&gt;Later dynamics can transport this stored entanglement across different partitions.&lt;/p&gt;

&lt;p&gt;This picture naturally explains why transport-dominated systems like MBL exhibit their largest observable entanglement growth at intermediate initial entanglement.&lt;/p&gt;




&lt;h2&gt;
  
  
  A new perspective on entanglement dynamics
&lt;/h2&gt;

&lt;p&gt;The broader message is that entanglement dynamics is not just about &lt;strong&gt;creating&lt;/strong&gt; quantum information.&lt;/p&gt;

&lt;p&gt;It is equally about &lt;strong&gt;processing, redistributing, and transporting&lt;/strong&gt; the quantum information that already exists.&lt;/p&gt;

&lt;p&gt;The traditional "start from product states" paradigm has taught us a great deal, but it captures only one half of the story.&lt;/p&gt;

&lt;p&gt;Distinguishing between &lt;strong&gt;entanglement generation&lt;/strong&gt; and &lt;strong&gt;entanglement transport&lt;/strong&gt; provides a more complete framework for understanding quantum dynamics across thermalizing systems, free fermions, many-body localization, and quantum circuits.&lt;/p&gt;

&lt;p&gt;Beyond offering a conceptual picture, this framework makes concrete predictions that can be tested on today's quantum simulation platforms, providing new experimental probes of nonequilibrium quantum dynamics.&lt;/p&gt;




&lt;h2&gt;
  
  
  Reference
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Entanglement growth from entangled states: A unified perspective on entanglement generation and transport&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Chun-Yue Zhang, Zi-Xiang Li, and Shi-Xin Zhang&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Physical Review Letters&lt;/em&gt; &lt;strong&gt;137&lt;/strong&gt;, 020404 (2026)&lt;/p&gt;

</description>
      <category>quantum</category>
    </item>
  </channel>
</rss>
