I recently released a new preprint exploring a simple question:
If two neural networks eventually behave almost the same, does that mean they also become internally the same?
The experiments suggest that — at least under the protocol I tested — the answer can be no.
Two neural networks can reach very similar predictive performance while still retaining measurable differences in their internal representations caused by their earlier training histories.
The paper is now available on arXiv:
Paper: https://arxiv.org/abs/2609.37836
Code and reproducibility repository:
https://github.com/Ertugrulmutlu/hysteresis-neural-networks
The basic idea
The experiments use MNIST and a small convolutional neural network.
Instead of training a model once on the full dataset, I split the learning history into stages.
Define:
- A: digits 0–4
- B: digits 5–9
- C: a balanced distribution containing digits 0–9
Two models start from identical initial weights.
The first model experiences:
A -> B -> C
The second experiences:
B -> A -> C
I call these:
SABC
SBAC
The important part is the final stage.
During C, both models receive the same training distribution, the same deterministic batch sequence, and the same checkpoint schedule.
So after their different histories, both networks are exposed to the same environment.
The question becomes:
Does common training erase the effects of training history?
Why accuracy alone is not enough
Suppose both networks eventually reach around the same accuracy.
It would be tempting to conclude that they have converged to essentially the same solution.
But prediction accuracy only tells us about behavior.
It does not tell us whether the internal features learned by the two networks are identical.
To study this, I compare internal representations using CKA-based similarity measurements.
The main representation-history score is:
H_repr = 1 - mean(CKA_conv2, CKA_fc1)
Higher values mean the two networks retain more representational difference.
The interesting case is therefore:
Behavioral difference -> small
while
Representational difference -> remains measurable
What happened?
In the main experiment, I ran 20 paired seeds.
The networks often became closely matched in predictive performance after common relaxation, while measurable representation differences remained.
This motivated a harder test:
What happens if both models continue training on the same distribution for much longer?
Stress test: 50,000 shared updates
I reran five paired seeds from initialization and extended the common-relaxation phase to:
10,000 updates
25,000 updates
50,000 updates
At 50,000 common updates, the mean representation-history score was:
H_repr = 0.1902
with a 95% bootstrap confidence interval of approximately:
[0.1611, 0.2193]
Meanwhile, predictive performance remained very close.
At 50,000 updates:
Mean SABC accuracy: 98.480%
Mean SBAC accuracy: 98.316%
Mean absolute accuracy gap: ~0.18 percentage points
So over the measured horizon, behavioral convergence did not imply representational convergence.
Importantly, this does not prove that the difference persists forever.
The correct interpretation is narrower:
The internal representation difference did not disappear within the tested 50,000-update horizon.
A second question: does the difference actually matter?
Representation geometry can differ without having any practical consequence.
So I added a fresh linear-probe experiment.
The CNN backbone is frozen.
A brand-new linear classifier is then trained on top of FC1 features.
Only the new linear head is trained.
The SABC and SBAC probes receive:
- identical initialization,
- identical examples,
- identical labels,
- identical minibatch ordering,
- identical optimization settings.
This lets us ask:
Is class information equally easy to extract from both representations?
An interesting low-data effect
I tested several amounts of labeled training data:
25 examples / class
50 examples / class
100 examples / class
500 examples / class
The mean SABC-minus-SBAC accuracy differences were:
25/class -> +1.005 percentage points
50/class -> +0.580 percentage points
100/class -> +0.288 percentage points
500/class -> +0.090 percentage points
The difference shrinks as more labeled data becomes available.
The 50-examples-per-class condition was especially interesting, so I repeated the probe using five different fresh-head seeds.
After averaging repeated probes within each of the 20 model seeds:
Mean paired difference: +0.496 pp
95% t CI: [+0.127, +0.866] pp
95% bootstrap CI: [+0.160, +0.833] pp
paired t-test: p = 0.0111
Wilcoxon: p = 0.0083
This suggests a useful distinction.
The two representations can provide similar high-data performance while differing in low-data readout efficiency.
In other words, training history may change how easy it is for a downstream classifier to extract useful information when supervision is limited.
What happens with more probe data?
At 500 examples per class, the mean difference was only:
+0.090 percentage points
I also ran an exploratory equivalence analysis using a ±0.5 percentage-point margin.
The 90% confidence interval was:
[-0.037, +0.217] pp
which falls completely inside the chosen equivalence interval.
So under that declared margin, the high-data linear-probe endpoints are practically equivalent.
That produces a two-regime picture:
Low labeled data:
training history affects linear readout efficiency.
High labeled data:
the difference becomes practically small.
Is this just caused by ReLU?
One possible mechanism is reduced plasticity caused by sparse ReLU activation patterns.
To explore this, I ran a matched-learning-rate comparison between:
ReLU
vs.
LeakyReLU
Both conditions were evaluated through 50,000 common-relaxation updates across five paired seeds.
At the final endpoint, the LeakyReLU-minus-ReLU difference in the representation-history score was approximately:
-0.0398
All five paired differences pointed in the same direction.
The bootstrap interval was approximately:
[-0.0772, -0.0176]
But with only five model seeds, classical paired significance tests were still inconclusive.
So I treat this as directional mechanism evidence, not a causal proof.
The result suggests that activation-mediated plasticity may contribute to the persistence of training-history effects.
Is the result only caused by different labels?
The original A/B split separates digits:
A = 0–4
B = 5–9
That raises an obvious concern.
Maybe the effect is simply caused by learning disjoint output classes in different orders.
To test this, I added a same-label rotated-MNIST control.
Both histories use the same class labels, but the input domains differ through symmetric rotations.
Across all five paired seeds, the models reached the predeclared behavioral-matching criterion while still retaining non-zero representation-history scores.
So the phenomenon is not limited to the original disjoint-label setup.
Reproducing the experiments
The full repository is available here:
https://github.com/Ertugrulmutlu/hysteresis-neural-networks
Clone it:
git clone https://github.com/Ertugrulmutlu/hysteresis-neural-networks.git
cd hysteresis-neural-networks
Create an environment:
python -m venv .venv
On Windows PowerShell:
.venv\Scripts\Activate.ps1
Install dependencies:
pip install -r requirements.txt
Run the basic AB / BA experiment
python -m src.train configs/mnist_v0_sab.yaml
python -m src.train configs/mnist_v0_sba.yaml
Then validate that the pair uses the intended matched initialization:
python -m src.validate_pair \
--run-a results/mnist_v0_SAB_seed1337_normnone \
--run-b results/mnist_v0_SBA_seed1337_normnone \
--json-out results/pair_validation.json
Run representation analysis
For example, CKA:
python -m src.analysis.cka \
--run-sab results/mnist_v0_SAB_seed1337_normnone \
--run-sba results/mnist_v0_SBA_seed1337_normnone \
--samples-per-class 200 \
--outdir plots/part2_seed1337/cka
The repository also includes analyses for:
training curves
weight trajectories
CKA
weight interpolation
activation health
common relaxation
paired multi-seed aggregation
linear probes
long-horizon relaxation
activation-function controls
rotated-MNIST controls
A note on reproducibility
I tried to make the pairwise experimental design explicit.
The paired runs control factors such as:
initialization
random seeds
evaluation data
common-relaxation batch order
checkpoint schedule
probe initialization
probe minibatch order
Run metadata is stored alongside the generated experiment artifacts.
The final paper-facing outputs are packaged under:
paper_artifacts/
while the underlying experiment outputs remain the source of truth.
What I think is the interesting part
The most interesting result to me is not simply that two networks end up with different weights.
That is expected.
The more interesting observation is the combination:
same initialization
+ different training histories
+ long exposure to the same later distribution
+ almost identical predictive performance
+ persistently different internal representations
And the linear-probe experiments suggest that those differences can sometimes affect how efficiently information is extracted downstream.
This creates an interesting distinction between:
behavioral convergence
and
representational convergence
They are not necessarily the same thing.
What this does NOT show
There are several important limitations.
This work does not establish that:
- all neural networks exhibit hysteresis,
- training-history effects persist permanently,
- the measured representation difference has a single causal mechanism,
- similar accuracy implies functionally identical networks,
- the result automatically generalizes beyond the tested architectures and datasets.
The experiments currently focus mainly on a small CNN and MNIST-derived protocols.
So I view the results as evidence of persistent training-history dependence under controlled experimental conditions, rather than a universal theorem about neural networks.
Where to go next
There are several natural extensions:
- wider or deeper architectures,
- larger datasets,
- transformers,
- continual-learning benchmarks,
- longer common-relaxation horizons,
- interventions targeting plasticity,
- representation alignment methods,
- studying whether history dependence affects adaptation to entirely new downstream tasks.
One question I find particularly interesting is whether similar effects appear in large models that are repeatedly fine-tuned or continually updated.
If two systems behave similarly today, how much of their different histories is still encoded internally?
That seems increasingly relevant as models become continuously trained, adapted, and deployed.
Paper and code
📄 Paper
https://arxiv.org/abs/2609.37836
💻 Code + reproducibility artifacts
https://github.com/Ertugrulmutlu/hysteresis-neural-networks
If you reproduce the experiments, find an issue, or want to test the idea on another architecture, feel free to open an issue or discussion on GitHub.
Citation
@article{mutlu2026behavioral,
title={Behavioral Convergence Without Representational Convergence: Persistent Training-History Dependence in Neural Networks},
author={Mutlu, Ertuğrul},
journal={arXiv preprint arXiv:2609.37836},
year={2026}
}
Top comments (0)