DEV Community

Cover image for Behavioral Convergence Without Representational Convergence: Reproducing Training-History Dependence in Neural Networks
Ertugrul
Ertugrul

Posted on

Behavioral Convergence Without Representational Convergence: Reproducing Training-History Dependence in Neural Networks

I recently released a new preprint exploring a simple question:

If two neural networks eventually behave almost the same, does that mean they also become internally the same?

The experiments suggest that — at least under the protocol I tested — the answer can be no.

Two neural networks can reach very similar predictive performance while still retaining measurable differences in their internal representations caused by their earlier training histories.

The paper is now available on arXiv:

Paper: https://arxiv.org/abs/2609.37836

Code and reproducibility repository:

https://github.com/Ertugrulmutlu/hysteresis-neural-networks


The basic idea

The experiments use MNIST and a small convolutional neural network.

Instead of training a model once on the full dataset, I split the learning history into stages.

Define:

  • A: digits 0–4
  • B: digits 5–9
  • C: a balanced distribution containing digits 0–9

Two models start from identical initial weights.

The first model experiences:

A -> B -> C
Enter fullscreen mode Exit fullscreen mode

The second experiences:

B -> A -> C
Enter fullscreen mode Exit fullscreen mode

I call these:

SABC
SBAC
Enter fullscreen mode Exit fullscreen mode

The important part is the final stage.

During C, both models receive the same training distribution, the same deterministic batch sequence, and the same checkpoint schedule.

So after their different histories, both networks are exposed to the same environment.

The question becomes:

Does common training erase the effects of training history?


Why accuracy alone is not enough

Suppose both networks eventually reach around the same accuracy.

It would be tempting to conclude that they have converged to essentially the same solution.

But prediction accuracy only tells us about behavior.

It does not tell us whether the internal features learned by the two networks are identical.

To study this, I compare internal representations using CKA-based similarity measurements.

The main representation-history score is:

H_repr = 1 - mean(CKA_conv2, CKA_fc1)
Enter fullscreen mode Exit fullscreen mode

Higher values mean the two networks retain more representational difference.

The interesting case is therefore:

Behavioral difference -> small

while

Representational difference -> remains measurable
Enter fullscreen mode Exit fullscreen mode

What happened?

In the main experiment, I ran 20 paired seeds.

The networks often became closely matched in predictive performance after common relaxation, while measurable representation differences remained.

This motivated a harder test:

What happens if both models continue training on the same distribution for much longer?


Stress test: 50,000 shared updates

I reran five paired seeds from initialization and extended the common-relaxation phase to:

10,000 updates
25,000 updates
50,000 updates
Enter fullscreen mode Exit fullscreen mode

At 50,000 common updates, the mean representation-history score was:

H_repr = 0.1902
Enter fullscreen mode Exit fullscreen mode

with a 95% bootstrap confidence interval of approximately:

[0.1611, 0.2193]
Enter fullscreen mode Exit fullscreen mode

Meanwhile, predictive performance remained very close.

At 50,000 updates:

Mean SABC accuracy: 98.480%
Mean SBAC accuracy: 98.316%

Mean absolute accuracy gap: ~0.18 percentage points
Enter fullscreen mode Exit fullscreen mode

So over the measured horizon, behavioral convergence did not imply representational convergence.

Importantly, this does not prove that the difference persists forever.

The correct interpretation is narrower:

The internal representation difference did not disappear within the tested 50,000-update horizon.


A second question: does the difference actually matter?

Representation geometry can differ without having any practical consequence.

So I added a fresh linear-probe experiment.

The CNN backbone is frozen.

A brand-new linear classifier is then trained on top of FC1 features.

Only the new linear head is trained.

The SABC and SBAC probes receive:

  • identical initialization,
  • identical examples,
  • identical labels,
  • identical minibatch ordering,
  • identical optimization settings.

This lets us ask:

Is class information equally easy to extract from both representations?


An interesting low-data effect

I tested several amounts of labeled training data:

25 examples / class
50 examples / class
100 examples / class
500 examples / class
Enter fullscreen mode Exit fullscreen mode

The mean SABC-minus-SBAC accuracy differences were:

25/class  -> +1.005 percentage points
50/class  -> +0.580 percentage points
100/class -> +0.288 percentage points
500/class -> +0.090 percentage points
Enter fullscreen mode Exit fullscreen mode

The difference shrinks as more labeled data becomes available.

The 50-examples-per-class condition was especially interesting, so I repeated the probe using five different fresh-head seeds.

After averaging repeated probes within each of the 20 model seeds:

Mean paired difference: +0.496 pp
95% t CI:              [+0.127, +0.866] pp
95% bootstrap CI:      [+0.160, +0.833] pp
paired t-test:         p = 0.0111
Wilcoxon:              p = 0.0083
Enter fullscreen mode Exit fullscreen mode

This suggests a useful distinction.

The two representations can provide similar high-data performance while differing in low-data readout efficiency.

In other words, training history may change how easy it is for a downstream classifier to extract useful information when supervision is limited.


What happens with more probe data?

At 500 examples per class, the mean difference was only:

+0.090 percentage points
Enter fullscreen mode Exit fullscreen mode

I also ran an exploratory equivalence analysis using a ±0.5 percentage-point margin.

The 90% confidence interval was:

[-0.037, +0.217] pp
Enter fullscreen mode Exit fullscreen mode

which falls completely inside the chosen equivalence interval.

So under that declared margin, the high-data linear-probe endpoints are practically equivalent.

That produces a two-regime picture:

Low labeled data:
training history affects linear readout efficiency.

High labeled data:
the difference becomes practically small.
Enter fullscreen mode Exit fullscreen mode

Is this just caused by ReLU?

One possible mechanism is reduced plasticity caused by sparse ReLU activation patterns.

To explore this, I ran a matched-learning-rate comparison between:

ReLU
vs.
LeakyReLU
Enter fullscreen mode Exit fullscreen mode

Both conditions were evaluated through 50,000 common-relaxation updates across five paired seeds.

At the final endpoint, the LeakyReLU-minus-ReLU difference in the representation-history score was approximately:

-0.0398
Enter fullscreen mode Exit fullscreen mode

All five paired differences pointed in the same direction.

The bootstrap interval was approximately:

[-0.0772, -0.0176]
Enter fullscreen mode Exit fullscreen mode

But with only five model seeds, classical paired significance tests were still inconclusive.

So I treat this as directional mechanism evidence, not a causal proof.

The result suggests that activation-mediated plasticity may contribute to the persistence of training-history effects.


Is the result only caused by different labels?

The original A/B split separates digits:

A = 0–4
B = 5–9
Enter fullscreen mode Exit fullscreen mode

That raises an obvious concern.

Maybe the effect is simply caused by learning disjoint output classes in different orders.

To test this, I added a same-label rotated-MNIST control.

Both histories use the same class labels, but the input domains differ through symmetric rotations.

Across all five paired seeds, the models reached the predeclared behavioral-matching criterion while still retaining non-zero representation-history scores.

So the phenomenon is not limited to the original disjoint-label setup.


Reproducing the experiments

The full repository is available here:

https://github.com/Ertugrulmutlu/hysteresis-neural-networks

Clone it:

git clone https://github.com/Ertugrulmutlu/hysteresis-neural-networks.git
cd hysteresis-neural-networks
Enter fullscreen mode Exit fullscreen mode

Create an environment:

python -m venv .venv
Enter fullscreen mode Exit fullscreen mode

On Windows PowerShell:

.venv\Scripts\Activate.ps1
Enter fullscreen mode Exit fullscreen mode

Install dependencies:

pip install -r requirements.txt
Enter fullscreen mode Exit fullscreen mode

Run the basic AB / BA experiment

python -m src.train configs/mnist_v0_sab.yaml
python -m src.train configs/mnist_v0_sba.yaml
Enter fullscreen mode Exit fullscreen mode

Then validate that the pair uses the intended matched initialization:

python -m src.validate_pair \
  --run-a results/mnist_v0_SAB_seed1337_normnone \
  --run-b results/mnist_v0_SBA_seed1337_normnone \
  --json-out results/pair_validation.json
Enter fullscreen mode Exit fullscreen mode

Run representation analysis

For example, CKA:

python -m src.analysis.cka \
  --run-sab results/mnist_v0_SAB_seed1337_normnone \
  --run-sba results/mnist_v0_SBA_seed1337_normnone \
  --samples-per-class 200 \
  --outdir plots/part2_seed1337/cka
Enter fullscreen mode Exit fullscreen mode

The repository also includes analyses for:

training curves
weight trajectories
CKA
weight interpolation
activation health
common relaxation
paired multi-seed aggregation
linear probes
long-horizon relaxation
activation-function controls
rotated-MNIST controls
Enter fullscreen mode Exit fullscreen mode

A note on reproducibility

I tried to make the pairwise experimental design explicit.

The paired runs control factors such as:

initialization
random seeds
evaluation data
common-relaxation batch order
checkpoint schedule
probe initialization
probe minibatch order
Enter fullscreen mode Exit fullscreen mode

Run metadata is stored alongside the generated experiment artifacts.

The final paper-facing outputs are packaged under:

paper_artifacts/
Enter fullscreen mode Exit fullscreen mode

while the underlying experiment outputs remain the source of truth.


What I think is the interesting part

The most interesting result to me is not simply that two networks end up with different weights.

That is expected.

The more interesting observation is the combination:

same initialization

+ different training histories

+ long exposure to the same later distribution

+ almost identical predictive performance

+ persistently different internal representations
Enter fullscreen mode Exit fullscreen mode

And the linear-probe experiments suggest that those differences can sometimes affect how efficiently information is extracted downstream.

This creates an interesting distinction between:

behavioral convergence

and

representational convergence
Enter fullscreen mode Exit fullscreen mode

They are not necessarily the same thing.


What this does NOT show

There are several important limitations.

This work does not establish that:

  • all neural networks exhibit hysteresis,
  • training-history effects persist permanently,
  • the measured representation difference has a single causal mechanism,
  • similar accuracy implies functionally identical networks,
  • the result automatically generalizes beyond the tested architectures and datasets.

The experiments currently focus mainly on a small CNN and MNIST-derived protocols.

So I view the results as evidence of persistent training-history dependence under controlled experimental conditions, rather than a universal theorem about neural networks.


Where to go next

There are several natural extensions:

  • wider or deeper architectures,
  • larger datasets,
  • transformers,
  • continual-learning benchmarks,
  • longer common-relaxation horizons,
  • interventions targeting plasticity,
  • representation alignment methods,
  • studying whether history dependence affects adaptation to entirely new downstream tasks.

One question I find particularly interesting is whether similar effects appear in large models that are repeatedly fine-tuned or continually updated.

If two systems behave similarly today, how much of their different histories is still encoded internally?

That seems increasingly relevant as models become continuously trained, adapted, and deployed.


Paper and code

📄 Paper

https://arxiv.org/abs/2609.37836

💻 Code + reproducibility artifacts

https://github.com/Ertugrulmutlu/hysteresis-neural-networks

If you reproduce the experiments, find an issue, or want to test the idea on another architecture, feel free to open an issue or discussion on GitHub.


Citation

@article{mutlu2026behavioral,
  title={Behavioral Convergence Without Representational Convergence: Persistent Training-History Dependence in Neural Networks},
  author={Mutlu, Ertuğrul},
  journal={arXiv preprint arXiv:2609.37836},
  year={2026}
}
Enter fullscreen mode Exit fullscreen mode

Top comments (0)