<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ertugrul</title>
    <description>The latest articles on DEV Community by Ertugrul (@ertugrulmutlu).</description>
    <link>https://dev.to/ertugrulmutlu</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1342871%2F80fb281c-da45-4841-b3a7-762c28f06416.jpg</url>
      <title>DEV Community: Ertugrul</title>
      <link>https://dev.to/ertugrulmutlu</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ertugrulmutlu"/>
    <language>en</language>
    <item>
      <title>Behavioral Convergence Without Representational Convergence: Reproducing Training-History Dependence in Neural Networks</title>
      <dc:creator>Ertugrul</dc:creator>
      <pubDate>Wed, 30 Sep 2026 14:34:04 +0000</pubDate>
      <link>https://dev.to/ertugrulmutlu/behavioral-convergence-without-representational-convergence-reproducing-training-history-45e6</link>
      <guid>https://dev.to/ertugrulmutlu/behavioral-convergence-without-representational-convergence-reproducing-training-history-45e6</guid>
      <description>&lt;p&gt;I recently released a new preprint exploring a simple question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If two neural networks eventually behave almost the same, does that mean they also become internally the same?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The experiments suggest that — at least under the protocol I tested — the answer can be &lt;strong&gt;no&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Two neural networks can reach very similar predictive performance while still retaining measurable differences in their internal representations caused by their earlier training histories.&lt;/p&gt;

&lt;p&gt;The paper is now available on arXiv:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Paper:&lt;/strong&gt; &lt;a href="https://arxiv.org/abs/2609.37836" rel="noopener noreferrer"&gt;https://arxiv.org/abs/2609.37836&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code and reproducibility repository:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;a href="https://github.com/Ertugrulmutlu/hysteresis-neural-networks" rel="noopener noreferrer"&gt;https://github.com/Ertugrulmutlu/hysteresis-neural-networks&lt;/a&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  The basic idea
&lt;/h2&gt;

&lt;p&gt;The experiments use MNIST and a small convolutional neural network.&lt;/p&gt;

&lt;p&gt;Instead of training a model once on the full dataset, I split the learning history into stages.&lt;/p&gt;

&lt;p&gt;Define:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A:&lt;/strong&gt; digits 0–4&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;B:&lt;/strong&gt; digits 5–9&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;C:&lt;/strong&gt; a balanced distribution containing digits 0–9&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Two models start from &lt;strong&gt;identical initial weights&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The first model experiences:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A -&amp;gt; B -&amp;gt; C
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The second experiences:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;B -&amp;gt; A -&amp;gt; C
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I call these:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SABC
SBAC
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important part is the final stage.&lt;/p&gt;

&lt;p&gt;During &lt;strong&gt;C&lt;/strong&gt;, both models receive the same training distribution, the same deterministic batch sequence, and the same checkpoint schedule.&lt;/p&gt;

&lt;p&gt;So after their different histories, both networks are exposed to the same environment.&lt;/p&gt;

&lt;p&gt;The question becomes:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Does common training erase the effects of training history?&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Why accuracy alone is not enough
&lt;/h2&gt;

&lt;p&gt;Suppose both networks eventually reach around the same accuracy.&lt;/p&gt;

&lt;p&gt;It would be tempting to conclude that they have converged to essentially the same solution.&lt;/p&gt;

&lt;p&gt;But prediction accuracy only tells us about &lt;strong&gt;behavior&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;It does not tell us whether the internal features learned by the two networks are identical.&lt;/p&gt;

&lt;p&gt;To study this, I compare internal representations using CKA-based similarity measurements.&lt;/p&gt;

&lt;p&gt;The main representation-history score is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;H_repr = 1 - mean(CKA_conv2, CKA_fc1)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Higher values mean the two networks retain more representational difference.&lt;/p&gt;

&lt;p&gt;The interesting case is therefore:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Behavioral difference -&amp;gt; small

while

Representational difference -&amp;gt; remains measurable
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  What happened?
&lt;/h2&gt;

&lt;p&gt;In the main experiment, I ran &lt;strong&gt;20 paired seeds&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The networks often became closely matched in predictive performance after common relaxation, while measurable representation differences remained.&lt;/p&gt;

&lt;p&gt;This motivated a harder test:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What happens if both models continue training on the same distribution for much longer?&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  Stress test: 50,000 shared updates
&lt;/h1&gt;

&lt;p&gt;I reran five paired seeds from initialization and extended the common-relaxation phase to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;10,000 updates
25,000 updates
50,000 updates
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At 50,000 common updates, the mean representation-history score was:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;H_repr = 0.1902
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;with a 95% bootstrap confidence interval of approximately:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[0.1611, 0.2193]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Meanwhile, predictive performance remained very close.&lt;/p&gt;

&lt;p&gt;At 50,000 updates:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Mean SABC accuracy: 98.480%
Mean SBAC accuracy: 98.316%

Mean absolute accuracy gap: ~0.18 percentage points
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So over the measured horizon, behavioral convergence did &lt;strong&gt;not&lt;/strong&gt; imply representational convergence.&lt;/p&gt;

&lt;p&gt;Importantly, this does &lt;strong&gt;not&lt;/strong&gt; prove that the difference persists forever.&lt;/p&gt;

&lt;p&gt;The correct interpretation is narrower:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The internal representation difference did not disappear within the tested 50,000-update horizon.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  A second question: does the difference actually matter?
&lt;/h1&gt;

&lt;p&gt;Representation geometry can differ without having any practical consequence.&lt;/p&gt;

&lt;p&gt;So I added a fresh linear-probe experiment.&lt;/p&gt;

&lt;p&gt;The CNN backbone is frozen.&lt;/p&gt;

&lt;p&gt;A brand-new linear classifier is then trained on top of FC1 features.&lt;/p&gt;

&lt;p&gt;Only the new linear head is trained.&lt;/p&gt;

&lt;p&gt;The SABC and SBAC probes receive:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;identical initialization,&lt;/li&gt;
&lt;li&gt;identical examples,&lt;/li&gt;
&lt;li&gt;identical labels,&lt;/li&gt;
&lt;li&gt;identical minibatch ordering,&lt;/li&gt;
&lt;li&gt;identical optimization settings.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This lets us ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Is class information equally easy to extract from both representations?&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  An interesting low-data effect
&lt;/h2&gt;

&lt;p&gt;I tested several amounts of labeled training data:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;25 examples / class
50 examples / class
100 examples / class
500 examples / class
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The mean SABC-minus-SBAC accuracy differences were:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;25/class  -&amp;gt; +1.005 percentage points
50/class  -&amp;gt; +0.580 percentage points
100/class -&amp;gt; +0.288 percentage points
500/class -&amp;gt; +0.090 percentage points
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The difference shrinks as more labeled data becomes available.&lt;/p&gt;

&lt;p&gt;The 50-examples-per-class condition was especially interesting, so I repeated the probe using five different fresh-head seeds.&lt;/p&gt;

&lt;p&gt;After averaging repeated probes within each of the 20 model seeds:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Mean paired difference: +0.496 pp
95% t CI:              [+0.127, +0.866] pp
95% bootstrap CI:      [+0.160, +0.833] pp
paired t-test:         p = 0.0111
Wilcoxon:              p = 0.0083
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This suggests a useful distinction.&lt;/p&gt;

&lt;p&gt;The two representations can provide similar high-data performance while differing in &lt;strong&gt;low-data readout efficiency&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;In other words, training history may change how easy it is for a downstream classifier to extract useful information when supervision is limited.&lt;/p&gt;




&lt;h1&gt;
  
  
  What happens with more probe data?
&lt;/h1&gt;

&lt;p&gt;At 500 examples per class, the mean difference was only:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;+0.090 percentage points
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I also ran an exploratory equivalence analysis using a ±0.5 percentage-point margin.&lt;/p&gt;

&lt;p&gt;The 90% confidence interval was:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[-0.037, +0.217] pp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;which falls completely inside the chosen equivalence interval.&lt;/p&gt;

&lt;p&gt;So under that declared margin, the high-data linear-probe endpoints are practically equivalent.&lt;/p&gt;

&lt;p&gt;That produces a two-regime picture:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Low labeled data:
training history affects linear readout efficiency.

High labeled data:
the difference becomes practically small.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  Is this just caused by ReLU?
&lt;/h1&gt;

&lt;p&gt;One possible mechanism is reduced plasticity caused by sparse ReLU activation patterns.&lt;/p&gt;

&lt;p&gt;To explore this, I ran a matched-learning-rate comparison between:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ReLU
vs.
LeakyReLU
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both conditions were evaluated through 50,000 common-relaxation updates across five paired seeds.&lt;/p&gt;

&lt;p&gt;At the final endpoint, the LeakyReLU-minus-ReLU difference in the representation-history score was approximately:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;-0.0398
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All five paired differences pointed in the same direction.&lt;/p&gt;

&lt;p&gt;The bootstrap interval was approximately:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[-0.0772, -0.0176]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But with only five model seeds, classical paired significance tests were still inconclusive.&lt;/p&gt;

&lt;p&gt;So I treat this as &lt;strong&gt;directional mechanism evidence&lt;/strong&gt;, not a causal proof.&lt;/p&gt;

&lt;p&gt;The result suggests that activation-mediated plasticity may contribute to the persistence of training-history effects.&lt;/p&gt;




&lt;h1&gt;
  
  
  Is the result only caused by different labels?
&lt;/h1&gt;

&lt;p&gt;The original A/B split separates digits:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A = 0–4
B = 5–9
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That raises an obvious concern.&lt;/p&gt;

&lt;p&gt;Maybe the effect is simply caused by learning disjoint output classes in different orders.&lt;/p&gt;

&lt;p&gt;To test this, I added a &lt;strong&gt;same-label rotated-MNIST control&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Both histories use the same class labels, but the input domains differ through symmetric rotations.&lt;/p&gt;

&lt;p&gt;Across all five paired seeds, the models reached the predeclared behavioral-matching criterion while still retaining non-zero representation-history scores.&lt;/p&gt;

&lt;p&gt;So the phenomenon is not limited to the original disjoint-label setup.&lt;/p&gt;




&lt;h1&gt;
  
  
  Reproducing the experiments
&lt;/h1&gt;

&lt;p&gt;The full repository is available here:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/Ertugrulmutlu/hysteresis-neural-networks" rel="noopener noreferrer"&gt;https://github.com/Ertugrulmutlu/hysteresis-neural-networks&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Clone it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/Ertugrulmutlu/hysteresis-neural-networks.git
&lt;span class="nb"&gt;cd &lt;/span&gt;hysteresis-neural-networks
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Create an environment:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python &lt;span class="nt"&gt;-m&lt;/span&gt; venv .venv
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On Windows PowerShell:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;venv&lt;/span&gt;&lt;span class="n"&gt;\Scripts\Activate.ps1&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Install dependencies:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; requirements.txt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Run the basic AB / BA experiment
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python &lt;span class="nt"&gt;-m&lt;/span&gt; src.train configs/mnist_v0_sab.yaml
python &lt;span class="nt"&gt;-m&lt;/span&gt; src.train configs/mnist_v0_sba.yaml
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then validate that the pair uses the intended matched initialization:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python &lt;span class="nt"&gt;-m&lt;/span&gt; src.validate_pair &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--run-a&lt;/span&gt; results/mnist_v0_SAB_seed1337_normnone &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--run-b&lt;/span&gt; results/mnist_v0_SBA_seed1337_normnone &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--json-out&lt;/span&gt; results/pair_validation.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Run representation analysis
&lt;/h2&gt;

&lt;p&gt;For example, CKA:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python &lt;span class="nt"&gt;-m&lt;/span&gt; src.analysis.cka &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--run-sab&lt;/span&gt; results/mnist_v0_SAB_seed1337_normnone &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--run-sba&lt;/span&gt; results/mnist_v0_SBA_seed1337_normnone &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--samples-per-class&lt;/span&gt; 200 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--outdir&lt;/span&gt; plots/part2_seed1337/cka
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The repository also includes analyses for:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;training curves
weight trajectories
CKA
weight interpolation
activation health
common relaxation
paired multi-seed aggregation
linear probes
long-horizon relaxation
activation-function controls
rotated-MNIST controls
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  A note on reproducibility
&lt;/h1&gt;

&lt;p&gt;I tried to make the pairwise experimental design explicit.&lt;/p&gt;

&lt;p&gt;The paired runs control factors such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;initialization
random seeds
evaluation data
common-relaxation batch order
checkpoint schedule
probe initialization
probe minibatch order
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run metadata is stored alongside the generated experiment artifacts.&lt;/p&gt;

&lt;p&gt;The final paper-facing outputs are packaged under:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;paper_artifacts/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;while the underlying experiment outputs remain the source of truth.&lt;/p&gt;




&lt;h1&gt;
  
  
  What I think is the interesting part
&lt;/h1&gt;

&lt;p&gt;The most interesting result to me is not simply that two networks end up with different weights.&lt;/p&gt;

&lt;p&gt;That is expected.&lt;/p&gt;

&lt;p&gt;The more interesting observation is the combination:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;same initialization

+ different training histories

+ long exposure to the same later distribution

+ almost identical predictive performance

+ persistently different internal representations
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And the linear-probe experiments suggest that those differences can sometimes affect how efficiently information is extracted downstream.&lt;/p&gt;

&lt;p&gt;This creates an interesting distinction between:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;behavioral convergence

and

representational convergence
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;They are not necessarily the same thing.&lt;/p&gt;




&lt;h1&gt;
  
  
  What this does NOT show
&lt;/h1&gt;

&lt;p&gt;There are several important limitations.&lt;/p&gt;

&lt;p&gt;This work does not establish that:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;all neural networks exhibit hysteresis,&lt;/li&gt;
&lt;li&gt;training-history effects persist permanently,&lt;/li&gt;
&lt;li&gt;the measured representation difference has a single causal mechanism,&lt;/li&gt;
&lt;li&gt;similar accuracy implies functionally identical networks,&lt;/li&gt;
&lt;li&gt;the result automatically generalizes beyond the tested architectures and datasets.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The experiments currently focus mainly on a small CNN and MNIST-derived protocols.&lt;/p&gt;

&lt;p&gt;So I view the results as evidence of &lt;strong&gt;persistent training-history dependence under controlled experimental conditions&lt;/strong&gt;, rather than a universal theorem about neural networks.&lt;/p&gt;




&lt;h1&gt;
  
  
  Where to go next
&lt;/h1&gt;

&lt;p&gt;There are several natural extensions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;wider or deeper architectures,&lt;/li&gt;
&lt;li&gt;larger datasets,&lt;/li&gt;
&lt;li&gt;transformers,&lt;/li&gt;
&lt;li&gt;continual-learning benchmarks,&lt;/li&gt;
&lt;li&gt;longer common-relaxation horizons,&lt;/li&gt;
&lt;li&gt;interventions targeting plasticity,&lt;/li&gt;
&lt;li&gt;representation alignment methods,&lt;/li&gt;
&lt;li&gt;studying whether history dependence affects adaptation to entirely new downstream tasks.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One question I find particularly interesting is whether similar effects appear in large models that are repeatedly fine-tuned or continually updated.&lt;/p&gt;

&lt;p&gt;If two systems behave similarly today, how much of their different histories is still encoded internally?&lt;/p&gt;

&lt;p&gt;That seems increasingly relevant as models become continuously trained, adapted, and deployed.&lt;/p&gt;




&lt;h2&gt;
  
  
  Paper and code
&lt;/h2&gt;

&lt;p&gt;📄 &lt;strong&gt;Paper&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/abs/2609.37836" rel="noopener noreferrer"&gt;https://arxiv.org/abs/2609.37836&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;💻 &lt;strong&gt;Code + reproducibility artifacts&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/Ertugrulmutlu/hysteresis-neural-networks" rel="noopener noreferrer"&gt;https://github.com/Ertugrulmutlu/hysteresis-neural-networks&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you reproduce the experiments, find an issue, or want to test the idea on another architecture, feel free to open an issue or discussion on GitHub.&lt;/p&gt;




&lt;h3&gt;
  
  
  Citation
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight bibtex"&gt;&lt;code&gt;&lt;span class="nc"&gt;@article&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;mutlu2026behavioral&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;title&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;{Behavioral Convergence Without Representational Convergence: Persistent Training-History Dependence in Neural Networks}&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;author&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;{Mutlu, Ertuğrul}&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;journal&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;{arXiv preprint arXiv:2609.37836}&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;year&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;{2026}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



</description>
      <category>ai</category>
      <category>computerscience</category>
      <category>machinelearning</category>
      <category>opensource</category>
    </item>
    <item>
      <title>What an Over-Engineered Parity Classifier Taught Me About Representation</title>
      <dc:creator>Ertugrul</dc:creator>
      <pubDate>Sat, 26 Sep 2026 16:23:53 +0000</pubDate>
      <link>https://dev.to/ertugrulmutlu/what-an-over-engineered-parity-classifier-taught-me-about-representation-40c8</link>
      <guid>https://dev.to/ertugrulmutlu/what-an-over-engineered-parity-classifier-taught-me-about-representation-40c8</guid>
      <description>&lt;h1&gt;
  
  
  I Revisited My Parity Paper — and Found the Representation Was the Real Story
&lt;/h1&gt;

&lt;p&gt;A while ago, I built a deliberately over-engineered classifier for one of the easiest problems in computer science:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is an integer odd or even?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;In binary, the answer is already sitting in the least significant bit.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;0&lt;/code&gt; means even.&lt;br&gt;&lt;br&gt;
&lt;code&gt;1&lt;/code&gt; means odd.&lt;/p&gt;

&lt;p&gt;No machine learning is needed.&lt;/p&gt;

&lt;p&gt;And yet I passed those binary representations through a wavelet transform, summarized the coefficients, clustered them with k-means, and tried to recover parity from the resulting feature space.&lt;/p&gt;

&lt;p&gt;The first version of the experiment looked interesting. It reported about &lt;strong&gt;69.67% accuracy&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;But when I came back to the project and started preparing a proper revision, I found something more important than the original result:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;the experiment itself needed to be rethought.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That ended up making the project much more interesting.&lt;/p&gt;




&lt;h2&gt;
  
  
  The two problems I found in the original version
&lt;/h2&gt;

&lt;p&gt;The first issue was &lt;strong&gt;label leakage&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I used parity labels to estimate whether each cluster was mostly odd or mostly even. But the evaluation was not cleanly separated from that calibration step.&lt;/p&gt;

&lt;p&gt;That meant information from the data being evaluated could influence the mapping from clusters to parity labels.&lt;/p&gt;

&lt;p&gt;The second issue was more conceptual.&lt;/p&gt;

&lt;p&gt;I had described the method as &lt;strong&gt;unsupervised&lt;/strong&gt; because k-means itself never receives parity labels.&lt;/p&gt;

&lt;p&gt;That is only partly true.&lt;/p&gt;

&lt;p&gt;The clustering step is unsupervised, but the final cluster-to-parity mapping uses labels. So the complete classifier is not fully unsupervised.&lt;/p&gt;

&lt;p&gt;Those two details changed how the result should be interpreted.&lt;/p&gt;

&lt;p&gt;Instead of trying to defend the original framing, I decided to rebuild the experiment around a stricter evaluation protocol.&lt;/p&gt;




&lt;h2&gt;
  
  
  Rebuilding the experiment
&lt;/h2&gt;

&lt;p&gt;The revised study uses all integers from &lt;code&gt;0&lt;/code&gt; to &lt;code&gt;10,000&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Each integer is converted into a fixed-width 32-bit binary signal.&lt;/p&gt;

&lt;p&gt;The primary pipeline is:&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Integer
   ↓
32-bit binary representation
   ↓
level-3 db2 wavelet transform
   ↓
mean absolute coefficient magnitude
   ↓
k-means per wavelet subband
   ↓
training-only cluster calibration
   ↓
odd / even prediction
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;The data is split into:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;6,000 training samples&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;2,000 validation samples&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;2,001 held-out test samples&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Representation and model choices are made using only training and validation data.&lt;/p&gt;

&lt;p&gt;Once the configuration is frozen, the model is re-fit on the combined train and validation sets and evaluated once on the held-out test set.&lt;/p&gt;

&lt;p&gt;The result:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;84.26% held-out test accuracy&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;with a 95% Wilson confidence interval of:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;82.60%–85.79%&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Across 20 different stratified train/test splits, the same representation achieved:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;84.20% ± 0.57%&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So the revised result is numerically stronger than the original one.&lt;/p&gt;

&lt;p&gt;But that is not the part I find most interesting.&lt;/p&gt;




&lt;h2&gt;
  
  
  The important result is what happens when the representation changes
&lt;/h2&gt;

&lt;p&gt;Parity is determined by one bit.&lt;/p&gt;

&lt;p&gt;So the cleanest test is simple:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;remove that bit.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When I mask the natural least significant bit and keep the rest of the wavelet pipeline unchanged, validation accuracy drops to:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;48.15%&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is essentially chance.&lt;/p&gt;

&lt;p&gt;This is important because it rules out the strongest interpretation of the model.&lt;/p&gt;

&lt;p&gt;The pipeline is not discovering the abstract arithmetic rule of parity.&lt;/p&gt;

&lt;p&gt;It is using information that is already present in the binary representation.&lt;/p&gt;

&lt;p&gt;The interesting question becomes:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why is that information so easy to recover in some representations and almost impossible to recover in others?&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  A tiny layout change can have a huge effect
&lt;/h2&gt;

&lt;p&gt;The standard representation uses left-zero padding.&lt;/p&gt;

&lt;p&gt;If I keep the same 32-bit signal but switch to right padding, validation accuracy drops to:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;65.20%&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The parity rule did not change.&lt;/p&gt;

&lt;p&gt;The bits did not contain less information.&lt;/p&gt;

&lt;p&gt;Only their &lt;strong&gt;position inside the signal&lt;/strong&gt; changed.&lt;/p&gt;

&lt;p&gt;That was the first strong hint that the wavelet transform was creating a geometry that depends heavily on spatial layout.&lt;/p&gt;




&lt;h2&gt;
  
  
  The coarse wavelet band carries almost all of the signal
&lt;/h2&gt;

&lt;p&gt;A level-3 wavelet decomposition produces one approximation band and three detail bands:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Subband&lt;/th&gt;
&lt;th&gt;Validation accuracy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A3&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;83.20%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;D3&lt;/td&gt;
&lt;td&gt;52.35%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;D2&lt;/td&gt;
&lt;td&gt;50.30%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;D1&lt;/td&gt;
&lt;td&gt;50.40%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This surprised me.&lt;/p&gt;

&lt;p&gt;Parity depends on a single bit, so I initially expected the fine-scale detail coefficients to matter most.&lt;/p&gt;

&lt;p&gt;Instead, the approximation band &lt;code&gt;A3&lt;/code&gt; contains almost the entire predictive signal.&lt;/p&gt;

&lt;p&gt;I then compared it with simpler coarse representations:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Representation&lt;/th&gt;
&lt;th&gt;Validation accuracy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;db2 A3&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;83.20%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Haar A3&lt;/td&gt;
&lt;td&gt;61.30%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Raw MAV&lt;/td&gt;
&lt;td&gt;61.30%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3-level average pooling&lt;/td&gt;
&lt;td&gt;61.30%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Triangular low-pass&lt;/td&gt;
&lt;td&gt;51.40%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;So this is not just an averaging effect.&lt;/p&gt;

&lt;p&gt;Something specific about the interaction between the db2 filters, downsampling, signal layout, and boundary handling makes the parity bit unusually accessible.&lt;/p&gt;




&lt;h2&gt;
  
  
  Then I moved the parity bit
&lt;/h2&gt;

&lt;p&gt;This was probably the most revealing experiment.&lt;/p&gt;

&lt;p&gt;I kept exactly the same 32 bits.&lt;/p&gt;

&lt;p&gt;I did not add information.&lt;/p&gt;

&lt;p&gt;I did not remove information.&lt;/p&gt;

&lt;p&gt;I only moved the parity-carrying bit to different positions in the signal.&lt;/p&gt;

&lt;p&gt;The result changed dramatically.&lt;/p&gt;

&lt;p&gt;At the natural position, validation accuracy is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;83.20%&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;At some positions, it falls much closer to chance.&lt;/p&gt;

&lt;p&gt;At the best tested position, it reaches:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;98.60%&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Nothing about the underlying parity information changed.&lt;/p&gt;

&lt;p&gt;Only its position changed.&lt;/p&gt;

&lt;p&gt;That makes the interpretation much clearer:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;the model is not learning a representation-independent rule.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It is exploiting a representation-dependent structure created by the transform.&lt;/p&gt;




&lt;h2&gt;
  
  
  Even boundary handling changes the result
&lt;/h2&gt;

&lt;p&gt;Wavelet transforms need a rule for what happens at the edges of a finite signal.&lt;/p&gt;

&lt;p&gt;I tested several boundary-extension modes while keeping the rest of the model fixed.&lt;/p&gt;

&lt;p&gt;The resulting validation accuracy ranged from:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;54.45% to 83.20%&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is a huge swing from what might look like a low-level implementation choice.&lt;/p&gt;

&lt;p&gt;In this experiment, boundary handling is not a minor detail.&lt;/p&gt;

&lt;p&gt;It is part of the mechanism.&lt;/p&gt;




&lt;h2&gt;
  
  
  Generalization exposes the weakness
&lt;/h2&gt;

&lt;p&gt;I also froze the model trained on &lt;code&gt;0–10,000&lt;/code&gt; and evaluated it on increasingly distant numerical ranges without recalibration.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Test range&lt;/th&gt;
&lt;th&gt;Accuracy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;10,001–20,000&lt;/td&gt;
&lt;td&gt;79.98%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;20,001–50,000&lt;/td&gt;
&lt;td&gt;71.51%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;50,001–100,000&lt;/td&gt;
&lt;td&gt;66.32%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;100,001–1,000,000&lt;/td&gt;
&lt;td&gt;59.69%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Performance steadily degrades as the magnitude distribution moves away from the training range.&lt;/p&gt;

&lt;p&gt;At first, that might suggest that larger integers are inherently harder.&lt;/p&gt;

&lt;p&gt;But when separate models are trained and tested inside fixed bit-length bands, accuracy remains roughly between &lt;strong&gt;78% and 88%&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;So the main problem is not magnitude itself.&lt;/p&gt;

&lt;p&gt;It is &lt;strong&gt;representation shift&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The geometry that works in one numerical regime does not stay stable in another.&lt;/p&gt;




&lt;h2&gt;
  
  
  More data does not solve it either
&lt;/h2&gt;

&lt;p&gt;On a wider &lt;code&gt;0–100,000&lt;/code&gt; distribution, I increased the training set from 500 examples all the way to 80,000.&lt;/p&gt;

&lt;p&gt;The performance ceiling barely moved.&lt;/p&gt;

&lt;p&gt;That suggests the bottleneck is not the amount of data.&lt;/p&gt;

&lt;p&gt;The bottleneck is the representation and the very simple clustering model operating on top of it.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I changed my mind about
&lt;/h2&gt;

&lt;p&gt;The first version of this project was mostly about a surprising classifier result.&lt;/p&gt;

&lt;p&gt;The revised version is not.&lt;/p&gt;

&lt;p&gt;The more interesting story is that the &lt;strong&gt;same symbolic information can become easy, difficult, or almost impossible to recover depending on how it is represented&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That changed how I think about the experiment.&lt;/p&gt;

&lt;p&gt;The question is no longer:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Can wavelets classify parity?”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Of course parity can be solved exactly with one bit.&lt;/p&gt;

&lt;p&gt;The better question is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“What does the representation make accessible to a simple model?”&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is a much more general machine-learning question.&lt;/p&gt;

&lt;p&gt;A model can perform well without learning the abstract rule we think it learned.&lt;/p&gt;

&lt;p&gt;Sometimes preprocessing creates a useful proxy.&lt;/p&gt;

&lt;p&gt;Sometimes position matters more than expected.&lt;/p&gt;

&lt;p&gt;Sometimes a boundary condition changes the geometry enough to alter the result completely.&lt;/p&gt;

&lt;p&gt;And sometimes the right experiment is not another accuracy benchmark, but an ablation that tells you &lt;strong&gt;where the accuracy came from&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why I am glad I revisited the paper
&lt;/h2&gt;

&lt;p&gt;Finding problems in an earlier experiment is uncomfortable.&lt;/p&gt;

&lt;p&gt;But I think revisiting it made the work much stronger.&lt;/p&gt;

&lt;p&gt;The revised version now:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;separates training, validation, and test data cleanly&lt;/li&gt;
&lt;li&gt;describes the supervised and unsupervised parts accurately&lt;/li&gt;
&lt;li&gt;includes repeated-split robustness checks&lt;/li&gt;
&lt;li&gt;tests out-of-distribution generalization&lt;/li&gt;
&lt;li&gt;isolates the role of the LSB&lt;/li&gt;
&lt;li&gt;studies subbands, padding, bit position, and boundary conditions&lt;/li&gt;
&lt;li&gt;includes frozen splits, predictions, metrics, and reproducibility artifacts&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The final claim is narrower than the original one.&lt;/p&gt;

&lt;p&gt;I am much more comfortable with it because of that.&lt;/p&gt;

&lt;p&gt;The wavelet pipeline does &lt;strong&gt;not&lt;/strong&gt; discover parity as an abstract arithmetic rule.&lt;/p&gt;

&lt;p&gt;Instead, it shows how a classical signal-processing transform can make symbolic information that is already present in the input more or less statistically accessible.&lt;/p&gt;

&lt;p&gt;For me, that ended up being the real result.&lt;/p&gt;




&lt;h2&gt;
  
  
  Reproducibility
&lt;/h2&gt;

&lt;p&gt;The code, frozen experiment artifacts, prediction outputs, and analysis results are available here:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GitHub:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;a href="https://github.com/Ertugrulmutlu/Using-Wavelets-and-Clustering-to-Predict-Odd-or-Even-Numbers" rel="noopener noreferrer"&gt;https://github.com/Ertugrulmutlu/Using-Wavelets-and-Clustering-to-Predict-Odd-or-Even-Numbers&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For the exact manuscript snapshot, use the Git tag:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;paper-v2&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;The revised arXiv version is scheduled to become public on September 29, 2026.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Paper:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;a href="https://arxiv.org/abs/2511.00071" rel="noopener noreferrer"&gt;https://arxiv.org/abs/2511.00071&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Final thought
&lt;/h2&gt;

&lt;p&gt;This project started as an unnecessarily complicated way to answer a one-bit question.&lt;/p&gt;

&lt;p&gt;The revision taught me something more useful than the original accuracy number:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Before asking what a model learned, ask what the representation made easy to learn.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is the part of this experiment I will probably remember.&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>python</category>
      <category>datascience</category>
      <category>signalprocessing</category>
    </item>
    <item>
      <title>PromptLedger v0.7 — Turning prompt evaluation into local regression gates</title>
      <dc:creator>Ertugrul</dc:creator>
      <pubDate>Mon, 13 Jul 2026 14:14:40 +0000</pubDate>
      <link>https://dev.to/ertugrulmutlu/promptledger-v07-turning-prompt-evaluation-into-local-regression-gates-4ano</link>
      <guid>https://dev.to/ertugrulmutlu/promptledger-v07-turning-prompt-evaluation-into-local-regression-gates-4ano</guid>
      <description>&lt;h2&gt;
  
  
  Devlog — Part 6
&lt;/h2&gt;

&lt;p&gt;PromptLedger v0.7 is out.&lt;/p&gt;

&lt;p&gt;The previous release made prompt history easier to inspect.&lt;/p&gt;

&lt;p&gt;This release makes prompt changes easier to evaluate.&lt;/p&gt;

&lt;p&gt;Until now, PromptLedger could answer questions such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What changed?&lt;/li&gt;
&lt;li&gt;Which version is in production?&lt;/li&gt;
&lt;li&gt;Which version was marked stable?&lt;/li&gt;
&lt;li&gt;Why was a new version created?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But one important question was still missing:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Did the new prompt actually perform better?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A text diff can show that a prompt changed.&lt;/p&gt;

&lt;p&gt;It cannot tell you whether accuracy improved, latency increased, cost became unacceptable, or an important behavior regressed.&lt;/p&gt;

&lt;p&gt;PromptLedger v0.7 introduces evaluation runs, metric comparisons, and policy-based regression gates.&lt;/p&gt;

&lt;p&gt;The workflow is now:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;version → diff → review → evaluate → gate → promote
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Evaluation runs
&lt;/h2&gt;

&lt;p&gt;PromptLedger can now store benchmark results produced by external tools.&lt;/p&gt;

&lt;p&gt;Each evaluation run is attached to a concrete prompt version and can include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;benchmark suite&lt;/li&gt;
&lt;li&gt;model&lt;/li&gt;
&lt;li&gt;dataset hash&lt;/li&gt;
&lt;li&gt;numeric metrics&lt;/li&gt;
&lt;li&gt;run metadata&lt;/li&gt;
&lt;li&gt;creation time&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"suite"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"support-v1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"test-model"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"metrics"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"accuracy"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.91&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"latency_ms"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;420&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"cost_usd"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.018&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"metadata"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"seed"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"sample_count"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A prompt version can have multiple evaluation runs.&lt;/p&gt;

&lt;p&gt;This is important because the same prompt may be tested:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;on multiple models&lt;/li&gt;
&lt;li&gt;against different datasets&lt;/li&gt;
&lt;li&gt;with different seeds&lt;/li&gt;
&lt;li&gt;across several benchmark suites&lt;/li&gt;
&lt;li&gt;multiple times as the surrounding system changes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Evaluation history is stored separately from prompt history, so repeated runs are preserved instead of overwriting each other.&lt;/p&gt;




&lt;h2&gt;
  
  
  Recording results
&lt;/h2&gt;

&lt;p&gt;PromptLedger does not run the benchmark itself.&lt;/p&gt;

&lt;p&gt;External tools produce the evaluation result, and PromptLedger records it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;promptledger &lt;span class="nb"&gt;eval &lt;/span&gt;record &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--id&lt;/span&gt; onboarding &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--ref&lt;/span&gt; staging &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--file&lt;/span&gt; evaluation-result.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The reference may be a concrete version or a label such as &lt;code&gt;prod&lt;/code&gt; or &lt;code&gt;staging&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Even when a moving label is used, the evaluation is stored against the concrete version the label currently points to.&lt;/p&gt;

&lt;p&gt;This keeps the historical record stable.&lt;/p&gt;




&lt;h2&gt;
  
  
  Comparing prompt versions
&lt;/h2&gt;

&lt;p&gt;Evaluation runs can be compared between two prompt versions or labels:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;promptledger &lt;span class="nb"&gt;eval &lt;/span&gt;compare &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--id&lt;/span&gt; onboarding &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--from&lt;/span&gt; prod &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--to&lt;/span&gt; staging &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--suite&lt;/span&gt; support-v1 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--model&lt;/span&gt; test-model
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Example output:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Evaluation comparison: onboarding
Suite: support-v1
Model: test-model

Metric          prod/v1     staging/v2    Delta
accuracy        0.84        0.91          +0.07
latency_ms      380         410           +30
cost_usd        0.016       0.018         +0.002
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;PromptLedger selects compatible runs deterministically.&lt;/p&gt;

&lt;p&gt;It does not silently compare results from unrelated benchmark suites or different models.&lt;/p&gt;

&lt;p&gt;A comparison is only useful when both sides describe the same experiment.&lt;/p&gt;




&lt;h2&gt;
  
  
  Regression gates
&lt;/h2&gt;

&lt;p&gt;Comparisons show what changed.&lt;/p&gt;

&lt;p&gt;Regression gates decide whether the change is acceptable.&lt;/p&gt;

&lt;p&gt;A gate policy defines:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;whether a metric should be higher or lower&lt;/li&gt;
&lt;li&gt;how much absolute regression is allowed&lt;/li&gt;
&lt;li&gt;how much percentage regression is allowed&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Example policy:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"suite"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"support-v1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"test-model"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"metrics"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"accuracy"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"direction"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"higher"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"max_regression"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.02&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"latency_ms"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"direction"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"lower"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"max_regression_percent"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;15&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"cost_usd"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"direction"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"lower"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"max_regression_percent"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run the gate with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;promptledger &lt;span class="nb"&gt;eval &lt;/span&gt;gate &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--id&lt;/span&gt; onboarding &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--from&lt;/span&gt; prod &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--to&lt;/span&gt; staging &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--policy&lt;/span&gt; promptledger-gate.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Example output:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;PASS accuracy: candidate improved by 0.07
PASS latency_ms: regression within allowed threshold
PASS cost_usd: regression within allowed threshold

Gate passed.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When a regression exceeds the policy limit, the command returns a non-zero exit code.&lt;/p&gt;

&lt;p&gt;That makes evaluation gates suitable for CI workflows.&lt;/p&gt;

&lt;p&gt;A prompt change can now fail a build in the same way as a failed test or performance regression.&lt;/p&gt;




&lt;h2&gt;
  
  
  Dashboard evaluation history
&lt;/h2&gt;

&lt;p&gt;The local dashboard has also been updated.&lt;/p&gt;

&lt;p&gt;Prompt details can now show:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;evaluation run history&lt;/li&gt;
&lt;li&gt;benchmark suite&lt;/li&gt;
&lt;li&gt;model&lt;/li&gt;
&lt;li&gt;dataset information&lt;/li&gt;
&lt;li&gt;metrics&lt;/li&gt;
&lt;li&gt;metadata&lt;/li&gt;
&lt;li&gt;metric differences between versions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This connects prompt history and evaluation history in the same interface.&lt;/p&gt;

&lt;p&gt;A developer can inspect what changed in the prompt and how the measured behavior changed alongside it.&lt;/p&gt;

&lt;p&gt;The dashboard remains local-first, and prompt content remains read-only.&lt;/p&gt;




&lt;h2&gt;
  
  
  Sequence-aware prompt diffs
&lt;/h2&gt;

&lt;p&gt;The dashboard comparison system was also improved.&lt;/p&gt;

&lt;p&gt;The previous implementation compared lines mainly by their array position.&lt;/p&gt;

&lt;p&gt;That works for simple replacements, but inserting one new line could make every following line appear changed.&lt;/p&gt;

&lt;p&gt;The new comparison uses a sequence-aware diff based on Python’s &lt;code&gt;difflib.SequenceMatcher&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;It now handles:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;inserted lines&lt;/li&gt;
&lt;li&gt;removed lines&lt;/li&gt;
&lt;li&gt;replaced lines&lt;/li&gt;
&lt;li&gt;unchanged sections&lt;/li&gt;
&lt;li&gt;aligned line numbers&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This produces a much more accurate view of how a prompt changed.&lt;/p&gt;




&lt;h2&gt;
  
  
  What changed in v0.7
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Dedicated evaluation history
&lt;/h3&gt;

&lt;p&gt;Benchmark results now live in their own evaluation-run records instead of being treated as static prompt metadata.&lt;/p&gt;

&lt;h3&gt;
  
  
  Version and label comparisons
&lt;/h3&gt;

&lt;p&gt;Evaluation results can be compared through concrete versions or release labels such as &lt;code&gt;prod&lt;/code&gt; and &lt;code&gt;staging&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Policy-based regression gates
&lt;/h3&gt;

&lt;p&gt;Teams can define acceptable quality, latency, and cost regressions through JSON policies.&lt;/p&gt;

&lt;h3&gt;
  
  
  CI-friendly exit codes
&lt;/h3&gt;

&lt;p&gt;Passing gates, detected regressions, invalid input, and operational failures return distinct exit codes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Dashboard evaluation visibility
&lt;/h3&gt;

&lt;p&gt;Evaluation runs and metric differences are now visible alongside prompt history.&lt;/p&gt;

&lt;h3&gt;
  
  
  Improved visual diffs
&lt;/h3&gt;

&lt;p&gt;Prompt comparisons now use a sequence-aware algorithm instead of simple line-position matching.&lt;/p&gt;

&lt;h3&gt;
  
  
  Non-destructive migration
&lt;/h3&gt;

&lt;p&gt;Existing PromptLedger databases are migrated to the new schema without removing prompt versions, labels, markers, or metadata.&lt;/p&gt;




&lt;h2&gt;
  
  
  Design boundary
&lt;/h2&gt;

&lt;p&gt;PromptLedger still does not execute prompts.&lt;/p&gt;

&lt;p&gt;It does not call:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;OpenAI&lt;/li&gt;
&lt;li&gt;Anthropic&lt;/li&gt;
&lt;li&gt;Gemini&lt;/li&gt;
&lt;li&gt;Ollama&lt;/li&gt;
&lt;li&gt;any other model provider&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It also does not include an automatic LLM judge.&lt;/p&gt;

&lt;p&gt;That separation is intentional.&lt;/p&gt;

&lt;p&gt;PromptLedger should remain:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;local-first&lt;/li&gt;
&lt;li&gt;provider-independent&lt;/li&gt;
&lt;li&gt;deterministic&lt;/li&gt;
&lt;li&gt;lightweight&lt;/li&gt;
&lt;li&gt;compatible with external benchmark systems&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Your evaluation framework runs the experiment.&lt;/p&gt;

&lt;p&gt;PromptLedger keeps the history, compares the results, and applies the release policy.&lt;/p&gt;




&lt;h2&gt;
  
  
  Testing
&lt;/h2&gt;

&lt;p&gt;The v0.7 release includes 119 passing tests covering:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;database migrations&lt;/li&gt;
&lt;li&gt;repeated evaluation runs&lt;/li&gt;
&lt;li&gt;label resolution&lt;/li&gt;
&lt;li&gt;malformed metrics&lt;/li&gt;
&lt;li&gt;non-finite values&lt;/li&gt;
&lt;li&gt;comparison compatibility&lt;/li&gt;
&lt;li&gt;deterministic ordering&lt;/li&gt;
&lt;li&gt;missing metrics&lt;/li&gt;
&lt;li&gt;zero baselines&lt;/li&gt;
&lt;li&gt;regression policies&lt;/li&gt;
&lt;li&gt;CLI exit codes&lt;/li&gt;
&lt;li&gt;dashboard endpoints&lt;/li&gt;
&lt;li&gt;sequence-aware diffs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The PowerShell file-based evaluation workflow was also tested end to end.&lt;/p&gt;




&lt;h2&gt;
  
  
  Installation
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;--upgrade&lt;/span&gt; promptledger
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Initialize a local database:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;promptledger init
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Launch the dashboard:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;promptledger dashboard
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;Prompt version control is useful, but version history alone cannot tell you whether a release is safe.&lt;/p&gt;

&lt;p&gt;A new prompt may look cleaner while producing worse answers.&lt;/p&gt;

&lt;p&gt;It may improve accuracy while increasing latency or cost.&lt;/p&gt;

&lt;p&gt;It may pass one benchmark while silently regressing another.&lt;/p&gt;

&lt;p&gt;PromptLedger v0.7 connects prompt changes to measurable results.&lt;/p&gt;

&lt;p&gt;The goal is not to build another evaluation framework.&lt;/p&gt;

&lt;p&gt;The goal is to provide the missing release layer between prompt experimentation and production:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;What changed?
How did performance change?
Is the regression acceptable?
Should this version be promoted?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;PromptLedger is moving from a tool that stores prompt history to a tool that helps control prompt releases.&lt;/p&gt;




&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;p&gt;PyPI: &lt;a href="https://pypi.org/project/promptledger/" rel="noopener noreferrer"&gt;PyPI&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;GitHub: &lt;a href="https://github.com/Ertugrulmutlu/promptledger" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;LinkedIn: &lt;a href="https://www.linkedin.com/in/ertugrul-mutlu/?locale=en" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Website: &lt;a href="https://ertugrulmutlu.github.io" rel="noopener noreferrer"&gt;Website&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>showdev</category>
      <category>testing</category>
    </item>
    <item>
      <title>One Gradient Spike, Six Batches to NaN: Debugging a Deterministic Neural-Network Training Failure</title>
      <dc:creator>Ertugrul</dc:creator>
      <pubDate>Mon, 13 Jul 2026 06:00:00 +0000</pubDate>
      <link>https://dev.to/ertugrulmutlu/one-gradient-spike-six-batches-to-nan-debugging-a-deterministic-neural-network-training-failure-4bdh</link>
      <guid>https://dev.to/ertugrulmutlu/one-gradient-spike-six-batches-to-nan-debugging-a-deterministic-neural-network-training-failure-4bdh</guid>
      <description>&lt;p&gt;The first visible NaN appeared in the final linear layer at batch 46.&lt;/p&gt;

&lt;p&gt;By then, the model was already effectively destroyed.&lt;/p&gt;

&lt;p&gt;Six batches earlier, every tensor was still finite. The loss was finite. The gradients were finite. The momentum buffers were finite. Nothing had crossed the clean boundary between a valid floating-point value and NaN.&lt;/p&gt;

&lt;p&gt;But the training state was already in a runaway regime.&lt;/p&gt;

&lt;p&gt;This is how an aggregate-analysis failure led me backward—from a corrupted checkpoint, to one reproducible batch, to the finite gradient spike that began the collapse.&lt;/p&gt;

&lt;h2&gt;
  
  
  The experiment
&lt;/h2&gt;

&lt;p&gt;I was running a sequential-training experiment on MNIST using a small CNN.&lt;/p&gt;

&lt;p&gt;The experiment compared two training histories:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;SABC:&lt;/strong&gt; digits 0–4, then digits 5–9, followed by a shared balanced 0–9 relaxation phase.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SBAC:&lt;/strong&gt; digits 5–9, then digits 0–4, followed by the same relaxation phase.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;As a mechanism-control experiment, I replaced the existing ReLU activation with LeakyReLU while keeping the rest of the protocol fixed.&lt;/p&gt;

&lt;p&gt;The setup was:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;SimpleCNN&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;no normalization&lt;/li&gt;
&lt;li&gt;LeakyReLU slope &lt;code&gt;0.01&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;SGD&lt;/li&gt;
&lt;li&gt;learning rate &lt;code&gt;0.05&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;momentum &lt;code&gt;0.9&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;no weight decay&lt;/li&gt;
&lt;li&gt;deterministic seeds: &lt;code&gt;101&lt;/code&gt;, &lt;code&gt;202&lt;/code&gt;, &lt;code&gt;303&lt;/code&gt;, &lt;code&gt;404&lt;/code&gt;, &lt;code&gt;505&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;two histories per seed&lt;/li&gt;
&lt;li&gt;ten total runs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Four of the ten runs became numerically invalid:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;SABC seed303&lt;/li&gt;
&lt;li&gt;SBAC seed101&lt;/li&gt;
&lt;li&gt;SBAC seed202&lt;/li&gt;
&lt;li&gt;SBAC seed303&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In all four cases, the first known bad checkpoint was epoch 11.&lt;/p&gt;

&lt;p&gt;Epoch 10 was the end of the first training domain. Epoch 11 was the first epoch after switching to the second domain.&lt;/p&gt;

&lt;p&gt;That timing was suspicious, but not enough to establish a cause. The instability could depend on the interaction between the activation function, learning rate, momentum, domain switch, model state, task order, and random seed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The validator found it first
&lt;/h2&gt;

&lt;p&gt;I did not discover the problem because the training loop printed &lt;code&gt;NaN&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;A later aggregate-analysis validator rejected two checkpoints that should have represented the same model state: the final sequential-training checkpoint and the step-zero checkpoint of the relaxation phase.&lt;/p&gt;

&lt;p&gt;At first, this looked like a serialization or bookkeeping issue.&lt;/p&gt;

&lt;p&gt;A direct checkpoint audit showed that the affected files contained NaN and Inf values.&lt;/p&gt;

&lt;p&gt;For three failed runs, all 421,642 parameters were NaN in the first bad checkpoint. In SABC seed303, the checkpoint contained 421,610 NaNs and 32 positive infinities.&lt;/p&gt;

&lt;p&gt;The validator had not caused the problem. It had prevented corrupted runs from entering the scientific analysis.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I initially got wrong
&lt;/h2&gt;

&lt;p&gt;The original runs had been launched in parallel on one GPU.&lt;/p&gt;

&lt;p&gt;My first suspicion was therefore environmental:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;GPU contention&lt;/li&gt;
&lt;li&gt;multiprocessing interference&lt;/li&gt;
&lt;li&gt;nondeterministic CUDA behavior&lt;/li&gt;
&lt;li&gt;a transient scheduling or hardware issue&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;To test that, I reran one known failing configuration—seed101-SBAC—alone with a detailed numerical tracer.&lt;/p&gt;

&lt;p&gt;It failed again at exactly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;epoch 11&lt;/li&gt;
&lt;li&gt;phase 2&lt;/li&gt;
&lt;li&gt;phase data A&lt;/li&gt;
&lt;li&gt;batch 46&lt;/li&gt;
&lt;li&gt;after the forward pass&lt;/li&gt;
&lt;li&gt;first failing tensor: &lt;code&gt;fc2_logits&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The isolated repeat also used the same input hash, label hash, label histogram, and non-finite counts.&lt;/p&gt;

&lt;p&gt;The failing logits contained:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;128 NaNs&lt;/li&gt;
&lt;li&gt;128 positive infinities&lt;/li&gt;
&lt;li&gt;384 negative infinities&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Reproducing the same failure in isolation, on the same batch and at the same network stage, makes a one-off GPU collision very unlikely.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually happened
&lt;/h2&gt;

&lt;p&gt;The first non-finite tensor appeared at batch 46.&lt;/p&gt;

&lt;p&gt;But the failure began earlier.&lt;/p&gt;

&lt;p&gt;At batch 39, the system still looked ordinary:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;loss: about &lt;code&gt;3.03&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;global gradient norm: about &lt;code&gt;9.11&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;parameter magnitudes still small&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;At batch 40, the trajectory changed sharply:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;loss jumped to about &lt;code&gt;44.59&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;global gradient norm jumped to about &lt;code&gt;483.74&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;maximum gradient reached about &lt;code&gt;266.7&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Everything was still finite.&lt;/p&gt;

&lt;p&gt;After the optimizer step, the largest momentum-buffer value rose from about &lt;code&gt;9.13&lt;/code&gt; to about &lt;code&gt;269.93&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The following batches did not recover.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Batch&lt;/th&gt;
&lt;th&gt;Loss&lt;/th&gt;
&lt;th&gt;Global gradient norm&lt;/th&gt;
&lt;th&gt;Max FC1 activation&lt;/th&gt;
&lt;th&gt;Interpretation&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;39&lt;/td&gt;
&lt;td&gt;&lt;code&gt;3.03&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;9.11&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;2.18&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Ordinary-looking state&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;40&lt;/td&gt;
&lt;td&gt;&lt;code&gt;44.59&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;483.74&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;17.10&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;First large finite gradient event&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;41&lt;/td&gt;
&lt;td&gt;&lt;code&gt;1.48 × 10^3&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;988&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;913&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Runaway growth becomes visible&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;43&lt;/td&gt;
&lt;td&gt;&lt;code&gt;1.04 × 10^5&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;7.82 × 10^4&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;5.92 × 10^4&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Strong amplification&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;44&lt;/td&gt;
&lt;td&gt;&lt;code&gt;4.71 × 10^8&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;6.25 × 10^7&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;1.02 × 10^6&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Extreme but finite&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;45&lt;/td&gt;
&lt;td&gt;&lt;code&gt;3.23 × 10^18&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;1.19 × 10^15&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;6.87 × 10^15&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Model effectively destroyed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;46&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;&lt;code&gt;1.97 × 10^36&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Final logits overflow float32&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Immediately before the failing forward pass:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;fc1.weight&lt;/code&gt; magnitude was around &lt;code&gt;10^12&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;fc2.weight&lt;/code&gt; magnitude was around &lt;code&gt;10^13&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;momentum buffers were around &lt;code&gt;10^14–10^15&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;the maximum FC1 activation was about &lt;code&gt;1.97 × 10^36&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The final linear layer then overflowed float32 and produced NaN and signed Inf values.&lt;/p&gt;

&lt;p&gt;The trace supports the following sequence:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;large but finite gradient event
→ rapid growth in momentum and parameter scale
→ rapidly growing activations
→ float32 overflow in the final linear layer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Momentum appears to have amplified the instability, but the trace does not establish what caused the initial gradient spike.&lt;/p&gt;

&lt;p&gt;Likewise, &lt;code&gt;fc2_logits&lt;/code&gt; was the first non-finite tensor, but it was the terminal symptom, not necessarily the origin of the failure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Finite is not healthy
&lt;/h2&gt;

&lt;p&gt;A standard safeguard is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;isfinite&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tensor&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;all&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="nf"&gt;stop_training&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That check is necessary, but late.&lt;/p&gt;

&lt;p&gt;At batch 45, the model had:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;loss around &lt;code&gt;10^18&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;gradient norm around &lt;code&gt;10^15&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;momentum buffers around &lt;code&gt;10^14&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;activations around &lt;code&gt;10^15&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Every value was still finite.&lt;/p&gt;

&lt;p&gt;A binary finite/non-finite check would classify the state as valid. In practice, the model was already beyond recovery.&lt;/p&gt;

&lt;p&gt;This suggests a more useful distinction:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;numerically finite&lt;/li&gt;
&lt;li&gt;numerically extreme&lt;/li&gt;
&lt;li&gt;numerically non-finite&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The transition from the first state to the second is where the most useful debugging information exists.&lt;/p&gt;

&lt;p&gt;By the time NaN appears, the original trigger may be several optimizer steps in the past.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the debugging workflow mattered
&lt;/h2&gt;

&lt;p&gt;The investigation followed a simple sequence:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Aggregate analysis failed.&lt;/li&gt;
&lt;li&gt;A pair validator caught a checkpoint inconsistency.&lt;/li&gt;
&lt;li&gt;A checkpoint audit found NaN and Inf values.&lt;/li&gt;
&lt;li&gt;A fail-fast numerical tracer was added.&lt;/li&gt;
&lt;li&gt;A known failure was reproduced in isolation.&lt;/li&gt;
&lt;li&gt;The first non-finite tensor was localized to one batch and one network stage.&lt;/li&gt;
&lt;li&gt;Earlier batches revealed the actual runaway dynamics.&lt;/li&gt;
&lt;li&gt;A finite control run showed that the tracer did not automatically cause failure.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That workflow is reusable.&lt;/p&gt;

&lt;p&gt;For numerical debugging, I would now monitor:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;global gradient norm&lt;/li&gt;
&lt;li&gt;maximum absolute gradient&lt;/li&gt;
&lt;li&gt;parameter norms&lt;/li&gt;
&lt;li&gt;momentum-buffer norms&lt;/li&gt;
&lt;li&gt;activation magnitudes&lt;/li&gt;
&lt;li&gt;model state before and after the optimizer step&lt;/li&gt;
&lt;li&gt;deterministic hashes of the input and labels&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The hashes were especially useful.&lt;/p&gt;

&lt;p&gt;“It failed at batch 46 again” is weaker evidence than:&lt;/p&gt;

&lt;p&gt;“It failed at batch 46 again on the exact same input and label tensors.”&lt;/p&gt;

&lt;h2&gt;
  
  
  What the evidence supports
&lt;/h2&gt;

&lt;p&gt;The careful conclusion is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Under this sequential-training protocol, LeakyReLU with SGD at learning rate 0.05 and momentum 0.9 showed deterministic, seed-dependent numerical instability in 4 of 10 runs.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The evidence does not show that LeakyReLU universally causes gradient explosions.&lt;/p&gt;

&lt;p&gt;It also does not show that momentum, the domain switch, the task order, or the final linear layer independently caused the failure.&lt;/p&gt;

&lt;p&gt;The instability may depend on an interaction between:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;activation function&lt;/li&gt;
&lt;li&gt;learning rate&lt;/li&gt;
&lt;li&gt;momentum&lt;/li&gt;
&lt;li&gt;domain switch&lt;/li&gt;
&lt;li&gt;model state&lt;/li&gt;
&lt;li&gt;random seed&lt;/li&gt;
&lt;li&gt;task order&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A control run, seed404-SBAC, completed the same requested epoch range without non-finite values.&lt;/p&gt;

&lt;p&gt;So the protocol was not universally unstable.&lt;/p&gt;

&lt;p&gt;But a 4/10 failure rate made this LeakyReLU branch unreliable for paired scientific analysis under the tested configuration.&lt;/p&gt;

&lt;p&gt;The existing ReLU results were not affected by this failure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The most useful artifact from this experiment was not the NaN checkpoint.&lt;/p&gt;

&lt;p&gt;It was the finite trace leading into it.&lt;/p&gt;

&lt;p&gt;At batch 39, the system looked ordinary.&lt;/p&gt;

&lt;p&gt;At batch 40, the gradient norm jumped from roughly &lt;code&gt;9&lt;/code&gt; to roughly &lt;code&gt;484&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;At batch 45, the model was still technically finite but already had gradients around &lt;code&gt;10^15&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;At batch 46, a hidden activation reached approximately &lt;code&gt;10^36&lt;/code&gt;, and the final linear layer overflowed float32.&lt;/p&gt;

&lt;p&gt;The main lesson is simple:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A finite tensor is not necessarily evidence of a healthy training state.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;NaN checks are essential, but they are late alarms.&lt;/p&gt;

&lt;p&gt;The final NaN showed where the computation stopped being representable.&lt;/p&gt;

&lt;p&gt;The six finite batches before it showed how the model got there.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>python</category>
      <category>learning</category>
    </item>
    <item>
      <title>OpenAnima v1 — Open-Source Desktop Overlay Engine for Windows</title>
      <dc:creator>Ertugrul</dc:creator>
      <pubDate>Fri, 15 May 2026 06:15:20 +0000</pubDate>
      <link>https://dev.to/ertugrulmutlu/openanima-v1-open-source-desktop-overlay-engine-for-windows-3m9c</link>
      <guid>https://dev.to/ertugrulmutlu/openanima-v1-open-source-desktop-overlay-engine-for-windows-3m9c</guid>
      <description>&lt;h1&gt;
  
  
  OpenAnima v1 — Open-Source Desktop Overlay Engine for Windows
&lt;/h1&gt;

&lt;p&gt;After months of building, experimenting, rewriting systems, debugging strange desktop issues, and testing different asset formats, OpenAnima v1 is finally out.&lt;/p&gt;

&lt;p&gt;OpenAnima is an open-source desktop overlay engine for Windows that lets you place animated assets directly onto your desktop.&lt;/p&gt;

&lt;p&gt;You can use:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;GIFs&lt;/li&gt;
&lt;li&gt;Static images&lt;/li&gt;
&lt;li&gt;Sprite strips&lt;/li&gt;
&lt;li&gt;Spritesheets&lt;/li&gt;
&lt;li&gt;Frame-folder animations&lt;/li&gt;
&lt;li&gt;RPG-style HUD elements&lt;/li&gt;
&lt;li&gt;Transparent animated assets&lt;/li&gt;
&lt;li&gt;Pixel-art characters&lt;/li&gt;
&lt;li&gt;Desktop companions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The goal of the project was simple:&lt;/p&gt;

&lt;p&gt;Create a lightweight desktop overlay system that feels flexible, customizable, and fun to experiment with.&lt;/p&gt;




&lt;h2&gt;
  
  
  What OpenAnima Can Do
&lt;/h2&gt;

&lt;p&gt;OpenAnima allows you to spawn movable overlay windows directly on top of your desktop.&lt;/p&gt;

&lt;p&gt;Each overlay can be:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;dragged freely&lt;/li&gt;
&lt;li&gt;resized&lt;/li&gt;
&lt;li&gt;locked in place&lt;/li&gt;
&lt;li&gt;made click-through&lt;/li&gt;
&lt;li&gt;set always-on-top&lt;/li&gt;
&lt;li&gt;adjusted for opacity&lt;/li&gt;
&lt;li&gt;animated with custom FPS settings&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The application also restores overlay states between sessions, so your desktop setup persists after restarting the app.&lt;/p&gt;




&lt;h2&gt;
  
  
  Supported Asset Types
&lt;/h2&gt;

&lt;p&gt;One of the biggest goals of the project was supporting multiple animation workflows instead of only GIFs.&lt;/p&gt;

&lt;p&gt;Currently supported:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;GIF animations&lt;/li&gt;
&lt;li&gt;Static images (.png, .jpg, .webp)&lt;/li&gt;
&lt;li&gt;Frame-folder animations&lt;/li&gt;
&lt;li&gt;Horizontal sprite strips&lt;/li&gt;
&lt;li&gt;Vertical sprite strips&lt;/li&gt;
&lt;li&gt;Spritesheets with metadata&lt;/li&gt;
&lt;li&gt;Basic HUD/UI overlay assets&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The engine also includes an asset analyzer system that tries to detect asset types automatically.&lt;/p&gt;




&lt;h2&gt;
  
  
  Built With
&lt;/h2&gt;

&lt;p&gt;OpenAnima is primarily built using:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Python&lt;/li&gt;
&lt;li&gt;PyQt6&lt;/li&gt;
&lt;li&gt;QMovie / QPixmap rendering systems&lt;/li&gt;
&lt;li&gt;Custom animation parsing logic&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A lot of time went into handling edge cases like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;hidden overlays&lt;/li&gt;
&lt;li&gt;broken configs&lt;/li&gt;
&lt;li&gt;off-screen windows&lt;/li&gt;
&lt;li&gt;click-through recovery&lt;/li&gt;
&lt;li&gt;corrupted animation states&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;v1 focuses heavily on stability and recovery tools.&lt;/p&gt;




&lt;h2&gt;
  
  
  Features Added During Development
&lt;/h2&gt;

&lt;p&gt;Some systems added throughout development:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Persistent overlay state saving&lt;/li&gt;
&lt;li&gt;Safer config recovery&lt;/li&gt;
&lt;li&gt;Recovery tools for invisible overlays&lt;/li&gt;
&lt;li&gt;Logging system&lt;/li&gt;
&lt;li&gt;Diagnostics panel&lt;/li&gt;
&lt;li&gt;Asset metadata support&lt;/li&gt;
&lt;li&gt;Import wizard&lt;/li&gt;
&lt;li&gt;Sprite animation handling&lt;/li&gt;
&lt;li&gt;Basic layered UI rendering&lt;/li&gt;
&lt;li&gt;System tray controls&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Why I Built It
&lt;/h2&gt;

&lt;p&gt;I always liked desktop customization tools, animated desktop companions, game HUD overlays, and lightweight desktop effects.&lt;/p&gt;

&lt;p&gt;Most existing solutions were either:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;too limited&lt;/li&gt;
&lt;li&gt;too complicated&lt;/li&gt;
&lt;li&gt;too specific&lt;/li&gt;
&lt;li&gt;abandoned&lt;/li&gt;
&lt;li&gt;or focused on only one asset type&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So I started experimenting with building my own system.&lt;/p&gt;

&lt;p&gt;The project slowly evolved from “just display GIFs on desktop” into a more general desktop overlay engine.&lt;/p&gt;




&lt;h2&gt;
  
  
  Future Plans
&lt;/h2&gt;

&lt;p&gt;Some ideas planned for future versions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Better transparent video support&lt;/li&gt;
&lt;li&gt;More advanced animation systems&lt;/li&gt;
&lt;li&gt;Improved UI tools&lt;/li&gt;
&lt;li&gt;Better performance optimizations&lt;/li&gt;
&lt;li&gt;Asset marketplace/import improvements&lt;/li&gt;
&lt;li&gt;More overlay interaction systems&lt;/li&gt;
&lt;li&gt;Possible 3D asset support in the future&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Download
&lt;/h2&gt;

&lt;p&gt;Website / Download:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;OpenAnima Website: &lt;a href="https://ertugrulmutlu.github.io/OpenAnima" rel="noopener noreferrer"&gt;https://ertugrulmutlu.github.io/OpenAnima/&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;GitHub Repository:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;OpenAnima GitHub Repository: &lt;a href="https://github.com/Ertugrulmutlu/OpenAnima" rel="noopener noreferrer"&gt;https://github.com/Ertugrulmutlu/OpenAnima&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;itch.io Page:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;OpenAnima itch.io Page: &lt;a href="https://ertugrulmutlu.itch.io/openanima" rel="noopener noreferrer"&gt;https://ertugrulmutlu.itch.io/openanima&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Instagram:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;OpenAnima Instagram: &lt;a href="https://www.instagram.com/openanimaengine" rel="noopener noreferrer"&gt;https://www.instagram.com/openanimaengine&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Final Thoughts
&lt;/h2&gt;

&lt;p&gt;OpenAnima started as a small side project and slowly became a much larger system than I originally expected.&lt;/p&gt;

&lt;p&gt;There is still a lot to improve, but releasing v1 feels like a huge milestone for the project.&lt;/p&gt;

&lt;p&gt;Feedback, ideas, bug reports, and experiments are always welcome.&lt;/p&gt;

&lt;p&gt;I’m excited to see what people create with it.&lt;/p&gt;

</description>
      <category>gamedev</category>
      <category>opensource</category>
      <category>showdev</category>
      <category>sideprojects</category>
    </item>
    <item>
      <title>PromptLedger v0.6 — Turning prompt history into a local workspace dashboard</title>
      <dc:creator>Ertugrul</dc:creator>
      <pubDate>Mon, 04 May 2026 16:23:46 +0000</pubDate>
      <link>https://dev.to/ertugrulmutlu/promptledger-v06-turning-prompt-history-into-a-local-workspace-dashboard-3cen</link>
      <guid>https://dev.to/ertugrulmutlu/promptledger-v06-turning-prompt-history-into-a-local-workspace-dashboard-3cen</guid>
      <description>&lt;h2&gt;
  
  
  Devlog — Part 5
&lt;/h2&gt;

&lt;p&gt;PromptLedger v0.6 is out.&lt;/p&gt;

&lt;p&gt;This release changes how PromptLedger feels to use.&lt;/p&gt;

&lt;p&gt;Until now, PromptLedger was primarily a terminal-first tool with a small read-only viewer. The core workflow already existed: store prompt versions, compare changes, label releases, mark important versions, and keep everything local in SQLite.&lt;/p&gt;

&lt;p&gt;That worked.&lt;/p&gt;

&lt;p&gt;But once a prompt library grows, a simple viewer stops being enough.&lt;/p&gt;

&lt;p&gt;Prompt iteration is not just about storing text. It is about navigating decisions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which version worked best?&lt;/li&gt;
&lt;li&gt;Which version became stable?&lt;/li&gt;
&lt;li&gt;What changed between versions?&lt;/li&gt;
&lt;li&gt;Which prompt belongs to which workflow?&lt;/li&gt;
&lt;li&gt;Which prompt should be reused?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those questions are easier to answer when the history is not only stored, but also visible and interactive.&lt;/p&gt;

&lt;p&gt;So v0.6 turns the old viewer into a local prompt workspace dashboard.&lt;/p&gt;




&lt;h2&gt;
  
  
  Workspace dashboard
&lt;/h2&gt;

&lt;p&gt;The dashboard now starts with a card-based workspace instead of a single prompt view.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fqpruubjud6h477k1pzua.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fqpruubjud6h477k1pzua.png" alt=" " width="800" height="469"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Each prompt appears as a card showing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;latest version&lt;/li&gt;
&lt;li&gt;collection and role&lt;/li&gt;
&lt;li&gt;markers such as stable or milestone&lt;/li&gt;
&lt;li&gt;short preview of the prompt&lt;/li&gt;
&lt;li&gt;last updated time&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This makes it easier to scan a prompt library and understand what exists without opening each item.&lt;/p&gt;




&lt;h2&gt;
  
  
  Prompt detail view
&lt;/h2&gt;

&lt;p&gt;Clicking a card opens a detailed view of the prompt.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fk1ukgab6e3ixuvfr39js.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fk1ukgab6e3ixuvfr39js.png" alt=" " width="800" height="619"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This view includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;full prompt text&lt;/li&gt;
&lt;li&gt;version metadata&lt;/li&gt;
&lt;li&gt;markers and labels&lt;/li&gt;
&lt;li&gt;version timeline&lt;/li&gt;
&lt;li&gt;side-by-side comparison&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The goal is to make it easy to understand how a prompt evolved over time.&lt;/p&gt;




&lt;h2&gt;
  
  
  What changed in v0.6
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Workspace instead of viewer
&lt;/h3&gt;

&lt;p&gt;The dashboard is no longer just a viewer. It is a workspace where prompts can be explored and organized visually.&lt;/p&gt;

&lt;h3&gt;
  
  
  Card-based interaction
&lt;/h3&gt;

&lt;p&gt;Prompts are now treated as objects instead of rows in a list. Cards provide quick context and actions.&lt;/p&gt;

&lt;h3&gt;
  
  
  Marker actions in the UI
&lt;/h3&gt;

&lt;p&gt;Stable and milestone markers can now be applied directly from the dashboard. These actions use the same underlying marker system as the CLI.&lt;/p&gt;

&lt;h3&gt;
  
  
  Compare workflow
&lt;/h3&gt;

&lt;p&gt;The compare view has been improved to clearly show differences between versions.&lt;/p&gt;

&lt;h3&gt;
  
  
  Usability improvements
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;keyboard shortcuts for faster navigation&lt;/li&gt;
&lt;li&gt;copy actions with feedback&lt;/li&gt;
&lt;li&gt;better empty states&lt;/li&gt;
&lt;li&gt;improved hover and selection behavior&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Design direction
&lt;/h2&gt;

&lt;p&gt;PromptLedger remains intentionally limited in scope.&lt;/p&gt;

&lt;p&gt;It is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;local-first&lt;/li&gt;
&lt;li&gt;SQLite-backed&lt;/li&gt;
&lt;li&gt;CLI-driven for write operations&lt;/li&gt;
&lt;li&gt;dashboard-driven for inspection and workflow&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It does not include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;cloud services&lt;/li&gt;
&lt;li&gt;telemetry&lt;/li&gt;
&lt;li&gt;external APIs&lt;/li&gt;
&lt;li&gt;AI features&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This boundary is important.&lt;/p&gt;

&lt;p&gt;The goal is not to build another platform, but to provide a reliable tool for working with prompt history.&lt;/p&gt;




&lt;h2&gt;
  
  
  Installation
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;--upgrade&lt;/span&gt; promptledger
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Run the dashboard
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;promptledger dashboard
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;v0.6 is a shift from storing prompt history to working with it.&lt;/p&gt;

&lt;p&gt;The dashboard is not meant to replace the CLI, but to complement it by making prompt iteration easier to inspect, compare, and organize.&lt;/p&gt;

&lt;p&gt;There is still a lot to improve, but this version establishes a clearer direction:&lt;/p&gt;

&lt;h2&gt;
  
  
  PromptLedger is becoming a tool for thinking about prompts, not just storing them.
&lt;/h2&gt;

&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;p&gt;PyPI: &lt;a href="https://pypi.org/project/promptledger/" rel="noopener noreferrer"&gt;PyPI&lt;/a&gt;&lt;br&gt;
GitHub: &lt;a href="https://github.com/Ertugrulmutlu/promptledger" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;&lt;br&gt;
LinkedIn: &lt;a href="https://www.linkedin.com/in/ertugrul-mutlu/?locale=en" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt;&lt;br&gt;
Website: &lt;a href="https://ertugrulmutlu.github.io" rel="noopener noreferrer"&gt;Website&lt;/a&gt;&lt;/p&gt;

</description>
      <category>promptengineering</category>
      <category>python</category>
      <category>opensource</category>
      <category>showdev</category>
    </item>
    <item>
      <title>OpenAnima v0.2 Preview: Turning the Windows Desktop into a Living Canvas</title>
      <dc:creator>Ertugrul</dc:creator>
      <pubDate>Fri, 01 May 2026 10:31:23 +0000</pubDate>
      <link>https://dev.to/ertugrulmutlu/openanima-v02-preview-turning-the-windows-desktop-into-a-living-canvas-11lh</link>
      <guid>https://dev.to/ertugrulmutlu/openanima-v02-preview-turning-the-windows-desktop-into-a-living-canvas-11lh</guid>
      <description>&lt;h1&gt;
  
  
  OpenAnima v0.2 Preview: Turning the Windows Desktop into a Living Canvas
&lt;/h1&gt;

&lt;p&gt;I recently published &lt;strong&gt;OpenAnima v0.2 Preview&lt;/strong&gt;, and this release is a big step for the project.&lt;/p&gt;

&lt;p&gt;OpenAnima started as a small experiment: what if I could place animated GIFs directly on my Windows desktop as movable overlay objects?&lt;/p&gt;

&lt;p&gt;That simple idea quickly became more interesting.&lt;/p&gt;

&lt;p&gt;Instead of only supporting GIFs, OpenAnima is now moving toward a more general goal:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;An open-source desktop asset overlay engine for Windows.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The idea is to make the desktop feel less static. OpenAnima lets you place animated assets, sprites, frame animations, HUD elements, and game-style visual assets directly on your desktop, then control how they behave.&lt;/p&gt;

&lt;p&gt;Website: &lt;a href="https://ertugrulmutlu.github.io/OpenAnima/" rel="noopener noreferrer"&gt;https://ertugrulmutlu.github.io/OpenAnima/&lt;/a&gt;&lt;br&gt;
Itch.io: &lt;a href="https://ertugrulmutlu.itch.io/openanima" rel="noopener noreferrer"&gt;https://ertugrulmutlu.itch.io/openanima&lt;/a&gt;&lt;br&gt;
GitHub: &lt;a href="https://github.com/Ertugrulmutlu/OpenAnima" rel="noopener noreferrer"&gt;https://github.com/Ertugrulmutlu/OpenAnima&lt;/a&gt;&lt;br&gt;
Release: &lt;a href="https://github.com/Ertugrulmutlu/OpenAnima/releases/tag/v0.2.0-preview" rel="noopener noreferrer"&gt;https://github.com/Ertugrulmutlu/OpenAnima/releases/tag/v0.2.0-preview&lt;/a&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  What OpenAnima does
&lt;/h2&gt;

&lt;p&gt;OpenAnima is a Windows desktop application that lets users add visual assets on top of their desktop.&lt;/p&gt;

&lt;p&gt;These assets can be moved around, configured, and used as lightweight desktop overlays.&lt;/p&gt;

&lt;p&gt;Some example use cases:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Desktop companions or animated mascots&lt;/li&gt;
&lt;li&gt;Game-style HUD elements&lt;/li&gt;
&lt;li&gt;Stream or recording overlays&lt;/li&gt;
&lt;li&gt;Sprite and animation preview experiments&lt;/li&gt;
&lt;li&gt;Ambient desktop widgets&lt;/li&gt;
&lt;li&gt;Weird little visual experiments&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The project is still early, but the direction is becoming clearer: OpenAnima is not just a GIF player. It is becoming a small visual layer above the desktop.&lt;/p&gt;


&lt;h2&gt;
  
  
  Why I built it
&lt;/h2&gt;

&lt;p&gt;I like projects that feel small at first, but slowly reveal a larger design space.&lt;/p&gt;

&lt;p&gt;At the beginning, OpenAnima was simply about putting GIFs on the desktop. That was already fun, but it also felt limited.&lt;/p&gt;

&lt;p&gt;Once I started thinking about game assets, HUD elements, sprites, and layered UI assets, the project became more like a desktop rendering playground.&lt;/p&gt;

&lt;p&gt;I wanted to explore questions like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Can desktop overlays be treated like reusable visual assets?&lt;/li&gt;
&lt;li&gt;Can a normal desktop become a small interactive canvas?&lt;/li&gt;
&lt;li&gt;Can game-style assets live outside a game engine?&lt;/li&gt;
&lt;li&gt;Can asset packs become something users can import, configure, and reuse?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is where the v0.2 release comes in.&lt;/p&gt;


&lt;h2&gt;
  
  
  What changed in v0.2
&lt;/h2&gt;

&lt;p&gt;The v0.2 preview expands OpenAnima from a GIF overlay tool into a more metadata-driven desktop asset engine.&lt;/p&gt;

&lt;p&gt;The biggest changes are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Generic asset analyzer and import wizard&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;asset.json&lt;/code&gt; metadata support&lt;/li&gt;
&lt;li&gt;Support for multiple asset formats&lt;/li&gt;
&lt;li&gt;Sprite strip setup workflow&lt;/li&gt;
&lt;li&gt;Spritesheet rendering with named animations&lt;/li&gt;
&lt;li&gt;Composite UI/HUD assets&lt;/li&gt;
&lt;li&gt;Runtime sliders for layered UI assets&lt;/li&gt;
&lt;li&gt;Improved Editor tab and Asset Setup dialog layouts&lt;/li&gt;
&lt;li&gt;Safer metadata validation and error handling&lt;/li&gt;
&lt;li&gt;Backward compatibility for existing GIF/static/frame-folder workflows&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This release is still a preview, but the foundation is much stronger than the first version.&lt;/p&gt;


&lt;h2&gt;
  
  
  Supported asset types
&lt;/h2&gt;

&lt;p&gt;OpenAnima v0.2 supports several asset types:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;GIF&lt;/li&gt;
&lt;li&gt;Static image&lt;/li&gt;
&lt;li&gt;Frame-folder animation&lt;/li&gt;
&lt;li&gt;Sprite strip&lt;/li&gt;
&lt;li&gt;Spritesheet&lt;/li&gt;
&lt;li&gt;Composite UI / HUD&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The first version was mostly focused on simple animated overlays. v0.2 starts building the system needed for more structured assets.&lt;/p&gt;


&lt;h2&gt;
  
  
  Metadata-driven assets
&lt;/h2&gt;

&lt;p&gt;One of the most important changes in v0.2 is &lt;code&gt;asset.json&lt;/code&gt; support.&lt;/p&gt;

&lt;p&gt;Instead of hardcoding every asset type, OpenAnima can now use metadata to understand how an asset should be loaded and rendered.&lt;/p&gt;

&lt;p&gt;A simplified example could look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"demo_hud"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"composite_ui"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"layers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"background"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"file"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"background.png"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"x"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"y"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"bar"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"file"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"bar.png"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"x"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;12&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"y"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This makes the project more flexible. Instead of treating every asset as just an image or GIF, OpenAnima can start understanding assets as structured objects.&lt;/p&gt;

&lt;p&gt;That opens the door for asset packs, richer UI overlays, and more advanced configuration later.&lt;/p&gt;




&lt;h2&gt;
  
  
  Sprite strips and spritesheets
&lt;/h2&gt;

&lt;p&gt;Another focus of v0.2 was better support for game-style assets.&lt;/p&gt;

&lt;p&gt;Sprite strips and spritesheets are common in game development, but they usually need some setup before they can be used correctly.&lt;/p&gt;

&lt;p&gt;OpenAnima now includes workflows for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Frame count&lt;/li&gt;
&lt;li&gt;Frame size&lt;/li&gt;
&lt;li&gt;Crop fields&lt;/li&gt;
&lt;li&gt;Preview grid&lt;/li&gt;
&lt;li&gt;Frame export&lt;/li&gt;
&lt;li&gt;Named animations for spritesheets&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is still not perfect, especially for unusual sprite layouts, but it is a good step toward making OpenAnima useful for more than simple GIF overlays.&lt;/p&gt;




&lt;h2&gt;
  
  
  Composite UI and HUD assets
&lt;/h2&gt;

&lt;p&gt;I also added support for composite UI/HUD assets.&lt;/p&gt;

&lt;p&gt;The idea is simple: an overlay does not have to be one image. It can be made from multiple layers.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A health bar background&lt;/li&gt;
&lt;li&gt;A fill layer&lt;/li&gt;
&lt;li&gt;A frame layer&lt;/li&gt;
&lt;li&gt;A text or icon layer&lt;/li&gt;
&lt;li&gt;Runtime sliders controlling values&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This makes it possible to experiment with game-like desktop HUDs.&lt;/p&gt;

&lt;p&gt;The current Composite UI editor is functional, but it is not meant to be a professional layout tool yet. It is more of a foundation for future versions.&lt;/p&gt;




&lt;h2&gt;
  
  
  The control panel
&lt;/h2&gt;

&lt;p&gt;OpenAnima has a small control panel with three main areas:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Library&lt;/li&gt;
&lt;li&gt;Active&lt;/li&gt;
&lt;li&gt;Editor&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The Library tab is used to import and manage assets.&lt;/p&gt;

&lt;p&gt;The Active tab is used to manage overlays currently placed on the desktop.&lt;/p&gt;

&lt;p&gt;The Editor tab is used to adjust selected assets with controls like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Scale&lt;/li&gt;
&lt;li&gt;Opacity&lt;/li&gt;
&lt;li&gt;Speed&lt;/li&gt;
&lt;li&gt;Always on top&lt;/li&gt;
&lt;li&gt;Click-through&lt;/li&gt;
&lt;li&gt;Locked&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The goal is to keep the interface practical and lightweight. I do not want OpenAnima to become a huge desktop suite. I want it to stay small, hackable, and focused.&lt;/p&gt;




&lt;h2&gt;
  
  
  Website and distribution
&lt;/h2&gt;

&lt;p&gt;For this release, I also created a small website for the project:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://ertugrulmutlu.github.io/OpenAnima/" rel="noopener noreferrer"&gt;https://ertugrulmutlu.github.io/OpenAnima/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The website explains the idea, shows the demo, links to the GitHub repository, and provides access to the Windows executable through GitHub Releases.&lt;/p&gt;

&lt;p&gt;I also published the project on itch.io as a free tool:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://ertugrulmutlu.itch.io/openanima" rel="noopener noreferrer"&gt;https://ertugrulmutlu.itch.io/openanima&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This was mainly to make the project feel more like a small product rather than only a repository.&lt;/p&gt;




&lt;h2&gt;
  
  
  Known limitations
&lt;/h2&gt;

&lt;p&gt;This is still a preview release, so some limitations are expected.&lt;/p&gt;

&lt;p&gt;Known issues include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Sprite strips may require manual frame count or frame size correction.&lt;/li&gt;
&lt;li&gt;Some sprite strips with unusual padding may need manual crop values.&lt;/li&gt;
&lt;li&gt;Spritesheets require metadata or setup through the import wizard.&lt;/li&gt;
&lt;li&gt;Composite UI assets may require manual layer alignment.&lt;/li&gt;
&lt;li&gt;The Composite UI editor is functional but not a full professional layout tool.&lt;/li&gt;
&lt;li&gt;3D model support is not included yet.&lt;/li&gt;
&lt;li&gt;Some unusual asset packs may still need manual &lt;code&gt;asset.json&lt;/code&gt; editing.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I prefer being clear about this because v0.2 is not a polished final product. It is a foundation release.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I want to explore next
&lt;/h2&gt;

&lt;p&gt;There are several directions I want to explore after v0.2:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Asset packs&lt;/li&gt;
&lt;li&gt;Better first-run experience&lt;/li&gt;
&lt;li&gt;More polished installer/distribution flow&lt;/li&gt;
&lt;li&gt;Better preview and import tools&lt;/li&gt;
&lt;li&gt;More reliable spritesheet workflows&lt;/li&gt;
&lt;li&gt;Simple 3D overlay experiments&lt;/li&gt;
&lt;li&gt;Linux experiments later&lt;/li&gt;
&lt;li&gt;Better documentation for custom assets&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The 3D direction is especially interesting, but I do not want to rush it before the 2D asset foundation is stable.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I learned
&lt;/h2&gt;

&lt;p&gt;This project reminded me that even small desktop tools can become surprisingly deep.&lt;/p&gt;

&lt;p&gt;At first, the problem looked simple:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Put an animated object on the desktop.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;But once I started supporting different asset types, the project became about importing, validating, describing, previewing, rendering, and controlling assets.&lt;/p&gt;

&lt;p&gt;That means the real challenge is not only drawing something on the desktop. The real challenge is creating a flexible system around desktop assets.&lt;/p&gt;

&lt;p&gt;OpenAnima v0.2 is my first serious step in that direction.&lt;/p&gt;




&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;p&gt;Website: &lt;a href="https://ertugrulmutlu.github.io/OpenAnima/" rel="noopener noreferrer"&gt;https://ertugrulmutlu.github.io/OpenAnima/&lt;/a&gt;&lt;br&gt;
GitHub: &lt;a href="https://github.com/Ertugrulmutlu/OpenAnima" rel="noopener noreferrer"&gt;https://github.com/Ertugrulmutlu/OpenAnima&lt;/a&gt;&lt;br&gt;
Release: &lt;a href="https://github.com/Ertugrulmutlu/OpenAnima/releases/tag/v0.2.0-preview" rel="noopener noreferrer"&gt;https://github.com/Ertugrulmutlu/OpenAnima/releases/tag/v0.2.0-preview&lt;/a&gt;&lt;br&gt;
itch.io: &lt;a href="https://ertugrulmutlu.itch.io/openanima" rel="noopener noreferrer"&gt;https://ertugrulmutlu.itch.io/openanima&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Feedback, bug reports, weird desktop overlay ideas, and asset workflow suggestions are very welcome.&lt;/p&gt;

&lt;p&gt;Thanks for reading.&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>showdev</category>
      <category>sideprojects</category>
      <category>ui</category>
    </item>
    <item>
      <title>PromptLedger v0.4 — Faster prompt logging, lightweight markers, and better prompt organization</title>
      <dc:creator>Ertugrul</dc:creator>
      <pubDate>Mon, 27 Apr 2026 15:19:18 +0000</pubDate>
      <link>https://dev.to/ertugrulmutlu/promptledger-v04-faster-prompt-logging-lightweight-markers-and-better-prompt-organization-2b2g</link>
      <guid>https://dev.to/ertugrulmutlu/promptledger-v04-faster-prompt-logging-lightweight-markers-and-better-prompt-organization-2b2g</guid>
      <description>&lt;p&gt;&lt;strong&gt;Devlog — Part 4&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;PromptLedger v0.4 is mostly about making prompt versioning easier to use repeatedly: faster logging, clearer organization, and small release signals for versions worth remembering.&lt;/p&gt;

&lt;p&gt;In the earlier parts of this series, PromptLedger started as a deliberately small local-first prompt version control tool. Then it gained labels, status semantics, better diffs, review workflows, semantic summaries, warnings, and Markdown export.&lt;/p&gt;

&lt;p&gt;Those additions made the history more useful after prompts had already been logged.&lt;/p&gt;

&lt;p&gt;Some of the direction for v0.4 also came from feedback and from watching where the workflow still felt a bit too manual. So before going into the details: thank you to everyone who tried the earlier versions, shared thoughts, pointed out rough edges, or simply asked practical questions about how the tool should behave during real prompt iteration.&lt;/p&gt;

&lt;p&gt;v0.4 focuses more on the moment before review: the actual day-to-day act of adding, organizing, and revisiting prompt versions.&lt;/p&gt;

&lt;p&gt;Because in practice, a prompt version control tool only works if logging a prompt is cheap enough that you actually keep doing it.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why this part was needed
&lt;/h2&gt;

&lt;p&gt;PromptLedger has always been intentionally limited in scope.&lt;/p&gt;

&lt;p&gt;It is SQLite-backed. It is terminal-first. It does not need a hosted backend. It does not try to execute prompts for you. It does not try to become an evaluation platform.&lt;/p&gt;

&lt;p&gt;The core job is still simple:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;store prompt versions&lt;/li&gt;
&lt;li&gt;compare them&lt;/li&gt;
&lt;li&gt;review changes&lt;/li&gt;
&lt;li&gt;organize prompt history&lt;/li&gt;
&lt;li&gt;make the history inspectable later&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But once I started using it more like a real prompt library, a small problem became obvious.&lt;/p&gt;

&lt;p&gt;Adding a prompt version was technically easy, but repetitive.&lt;/p&gt;

&lt;p&gt;During actual prompt iteration, the prompt text changes often, while the surrounding metadata usually stays the same. The same author. The same environment. The same tags. The same library grouping. The same role.&lt;/p&gt;

&lt;p&gt;Having to retype that metadata every time creates friction.&lt;/p&gt;

&lt;p&gt;And friction matters. If logging is annoying, the history becomes incomplete. If the history is incomplete, review becomes less useful. And if review becomes less useful, the tool stops doing its main job.&lt;/p&gt;

&lt;p&gt;So v0.4 is not a big architectural release.&lt;/p&gt;

&lt;p&gt;It is a workflow release.&lt;/p&gt;

&lt;p&gt;It makes repeated prompt logging faster, adds better organization primitives, and introduces lightweight markers for important versions.&lt;/p&gt;




&lt;h2&gt;
  
  
  Quick add: less typing during real iteration
&lt;/h2&gt;

&lt;p&gt;The main usability change in v0.4 is the new &lt;code&gt;add --quick&lt;/code&gt; workflow.&lt;/p&gt;

&lt;p&gt;A normal add can still be explicit:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;promptledger add &lt;span class="nt"&gt;--id&lt;/span&gt; onboarding &lt;span class="nt"&gt;--text&lt;/span&gt; &lt;span class="s2"&gt;"..."&lt;/span&gt; &lt;span class="nt"&gt;--collection&lt;/span&gt; support &lt;span class="nt"&gt;--role&lt;/span&gt; system
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is useful when creating a new prompt or when metadata should be stated clearly.&lt;/p&gt;

&lt;p&gt;But once a prompt already exists, most iterations do not need all metadata to be typed again. With &lt;code&gt;--quick&lt;/code&gt;, PromptLedger can reuse safe metadata defaults from the latest version of the same prompt id.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;promptledger add &lt;span class="nt"&gt;--id&lt;/span&gt; onboarding &lt;span class="nt"&gt;--text&lt;/span&gt; &lt;span class="s2"&gt;"..."&lt;/span&gt; &lt;span class="nt"&gt;--quick&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the latest &lt;code&gt;onboarding&lt;/code&gt; version already had metadata like author, tags, env, collection, and role, those values can be reused unless explicitly overridden.&lt;/p&gt;

&lt;p&gt;That means the common workflow becomes much smaller:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;edit the prompt&lt;/li&gt;
&lt;li&gt;add the new version&lt;/li&gt;
&lt;li&gt;keep moving&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This is not a flashy feature, but it changes the feel of the tool.&lt;/p&gt;

&lt;p&gt;Prompt logging becomes closer to a habit than a chore.&lt;/p&gt;

&lt;p&gt;That matters because PromptLedger is not useful because one perfect prompt was saved once. It is useful because a sequence of changes becomes reviewable later.&lt;/p&gt;

&lt;p&gt;Quick add is there to protect that sequence.&lt;/p&gt;




&lt;h2&gt;
  
  
  Collection and role are now first-class metadata
&lt;/h2&gt;

&lt;p&gt;v0.4 also adds two first-class metadata fields on prompt versions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;collection&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;role&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The reason is simple: once a prompt library grows, ids and tags are not always enough.&lt;/p&gt;

&lt;p&gt;A &lt;code&gt;collection&lt;/code&gt; gives prompts a lightweight grouping. It can represent a product area, a project, a use case, a customer support flow, an internal tool, or just a personal folder-like grouping.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;promptledger add &lt;span class="nt"&gt;--id&lt;/span&gt; onboarding &lt;span class="nt"&gt;--text&lt;/span&gt; &lt;span class="s2"&gt;"..."&lt;/span&gt; &lt;span class="nt"&gt;--collection&lt;/span&gt; support &lt;span class="nt"&gt;--role&lt;/span&gt; system
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This makes it easier to ask questions like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which prompts belong to the support collection?&lt;/li&gt;
&lt;li&gt;Which prompts are part of an onboarding workflow?&lt;/li&gt;
&lt;li&gt;Which versions were written for a specific environment or use case?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The second field, &lt;code&gt;role&lt;/code&gt;, is about what kind of prompt artifact this version represents.&lt;/p&gt;

&lt;p&gt;Built-in roles are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;system&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;user&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;template&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;modelfile&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;eval&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This matters because prompt libraries usually contain different kinds of artifacts.&lt;/p&gt;

&lt;p&gt;A system instruction is not the same thing as a reusable template. An evaluation prompt is not the same thing as a user-facing message. A model file prompt is not the same thing as an onboarding assistant instruction.&lt;/p&gt;

&lt;p&gt;Before v0.4, these differences could be represented with tags, but tags are free-form and tend to become messy over time.&lt;/p&gt;

&lt;p&gt;Making &lt;code&gt;role&lt;/code&gt; first-class gives PromptLedger a small amount of structure without turning it into a large framework.&lt;/p&gt;

&lt;p&gt;That is the balance I wanted here: enough organization to be useful, but not so much that the tool starts dictating how every prompt library must be designed.&lt;/p&gt;




&lt;h2&gt;
  
  
  Markers: small signals for important versions
&lt;/h2&gt;

&lt;p&gt;v0.4 introduces a new marker system for prompt versions.&lt;/p&gt;

&lt;p&gt;The core commands are:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;promptledger marker &lt;span class="nb"&gt;set&lt;/span&gt; &lt;span class="nt"&gt;--id&lt;/span&gt; onboarding &lt;span class="nt"&gt;--version&lt;/span&gt; 8 &lt;span class="nt"&gt;--name&lt;/span&gt; stable
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;promptledger marker remove &lt;span class="nt"&gt;--id&lt;/span&gt; onboarding &lt;span class="nt"&gt;--version&lt;/span&gt; 8 &lt;span class="nt"&gt;--name&lt;/span&gt; stable
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;promptledger marker list &lt;span class="nt"&gt;--id&lt;/span&gt; onboarding
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;promptledger marker show &lt;span class="nt"&gt;--id&lt;/span&gt; onboarding &lt;span class="nt"&gt;--version&lt;/span&gt; 8
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There are also convenience commands for the built-in markers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;promptledger stable &lt;span class="nt"&gt;--id&lt;/span&gt; onboarding
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;promptledger milestone &lt;span class="nt"&gt;--id&lt;/span&gt; onboarding
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The built-in markers are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;stable&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;milestone&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Markers are intentionally lighter than labels.&lt;/p&gt;

&lt;p&gt;That distinction is important.&lt;/p&gt;

&lt;p&gt;Labels in PromptLedger are useful when you want release-like semantics or named pointers. They can represent states such as a current production prompt, a reviewed version, or a version used in a specific workflow.&lt;/p&gt;

&lt;p&gt;Markers are smaller than that.&lt;/p&gt;

&lt;p&gt;A marker says: this version is worth noticing.&lt;/p&gt;

&lt;p&gt;Maybe it was the first version that worked well. Maybe it was a milestone during a rewrite. Maybe it was stable enough to revisit later. Maybe it is just a checkpoint that should not get lost in the version list.&lt;/p&gt;

&lt;p&gt;Not every important prompt version needs full label semantics.&lt;/p&gt;

&lt;p&gt;Sometimes you only need a small flag attached directly to a version.&lt;/p&gt;

&lt;p&gt;That is what markers are for.&lt;/p&gt;

&lt;p&gt;They are deliberately simple. They do not try to encode deployment state, environment ownership, evaluation results, or production routing. They are just lightweight release signals inside the local history.&lt;/p&gt;




&lt;h2&gt;
  
  
  Search, list, and show became more useful
&lt;/h2&gt;

&lt;p&gt;Organization only helps if the tools expose it.&lt;/p&gt;

&lt;p&gt;So v0.4 updates &lt;code&gt;list&lt;/code&gt;, &lt;code&gt;show&lt;/code&gt;, and &lt;code&gt;search&lt;/code&gt; to surface and use collection, role, and marker information.&lt;/p&gt;

&lt;p&gt;For example, searching by metadata becomes more natural:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;promptledger search &lt;span class="nt"&gt;--collection&lt;/span&gt; support &lt;span class="nt"&gt;--role&lt;/span&gt; system
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Search can also work with an empty &lt;code&gt;--contains&lt;/code&gt;, which means metadata-only filtering is now possible.&lt;/p&gt;

&lt;p&gt;That sounds small, but it changes how PromptLedger can be used.&lt;/p&gt;

&lt;p&gt;Before, search was mostly about finding prompt text. Now it can also be used to navigate the prompt library itself.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;show all system prompts in a collection&lt;/li&gt;
&lt;li&gt;find evaluation prompts across a project&lt;/li&gt;
&lt;li&gt;inspect versions marked as stable&lt;/li&gt;
&lt;li&gt;separate templates from user prompts&lt;/li&gt;
&lt;li&gt;review one slice of the library without relying on text search&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This moves PromptLedger a little further from “prompt version storage” toward “prompt library workflow”.&lt;/p&gt;

&lt;p&gt;Still local. Still small. Still inspectable.&lt;/p&gt;

&lt;p&gt;But more practical once the number of prompt artifacts grows.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Streamlit UI is still read-only
&lt;/h2&gt;

&lt;p&gt;PromptLedger also has a small Streamlit UI for inspecting prompt history.&lt;/p&gt;

&lt;p&gt;In v0.4, the UI now surfaces collection, role, and markers in timeline, detail, and comparison views.&lt;/p&gt;

&lt;p&gt;That makes the UI more useful when browsing a prompt library. You can see not only how a prompt changed, but also what kind of prompt it is, which collection it belongs to, and whether a version was marked as stable or as a milestone.&lt;/p&gt;

&lt;p&gt;The important part: the UI is still read-only.&lt;/p&gt;

&lt;p&gt;That is intentional.&lt;/p&gt;

&lt;p&gt;For now, PromptLedger keeps editing and logging in the CLI, while the UI remains focused on inspection. This keeps the implementation smaller and avoids turning the viewer into a second source of write behavior.&lt;/p&gt;

&lt;p&gt;The terminal is where versions are created.&lt;/p&gt;

&lt;p&gt;The UI is where history can be reviewed more comfortably.&lt;/p&gt;

&lt;p&gt;That split still feels right for the project.&lt;/p&gt;




&lt;h2&gt;
  
  
  Database changes stayed local and simple
&lt;/h2&gt;

&lt;p&gt;v0.4 moves the schema version from 3 to 5.&lt;/p&gt;

&lt;p&gt;The main database changes are straightforward:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a new &lt;code&gt;markers&lt;/code&gt; table&lt;/li&gt;
&lt;li&gt;new &lt;code&gt;collection&lt;/code&gt; and &lt;code&gt;role&lt;/code&gt; columns on &lt;code&gt;prompt_versions&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;migrations for the local SQLite database&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There is no remote migration service. No hosted registry. No account state. No backend coordination.&lt;/p&gt;

&lt;p&gt;The migration story stayed aligned with the rest of the project: local, inspectable, and boring in a good way.&lt;/p&gt;

&lt;p&gt;That matters because PromptLedger should remain easy to understand. A small SQLite-backed tool should not require infrastructure thinking just to store prompt history.&lt;/p&gt;




&lt;h2&gt;
  
  
  Tests expanded around the new workflows
&lt;/h2&gt;

&lt;p&gt;v0.4 also expands the test coverage substantially around the areas that changed most:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;add workflows&lt;/li&gt;
&lt;li&gt;quick add behavior&lt;/li&gt;
&lt;li&gt;marker commands&lt;/li&gt;
&lt;li&gt;search and metadata filtering&lt;/li&gt;
&lt;li&gt;related list/show behavior&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is not the kind of testing expansion that makes for a dramatic release note, but it is important for this project.&lt;/p&gt;

&lt;p&gt;PromptLedger deals with history. If history is stored incorrectly, the tool loses trust quickly.&lt;/p&gt;

&lt;p&gt;So the goal is practical correctness:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;metadata should be reused only when expected&lt;/li&gt;
&lt;li&gt;explicit overrides should still win&lt;/li&gt;
&lt;li&gt;markers should attach to the right versions&lt;/li&gt;
&lt;li&gt;search should return the intended slice of the prompt library&lt;/li&gt;
&lt;li&gt;existing workflows should not become harder to use&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For a local-first developer tool, confidence matters more than feature count.&lt;/p&gt;




&lt;h2&gt;
  
  
  Design tradeoffs
&lt;/h2&gt;

&lt;p&gt;The main tradeoff in v0.4 was deciding how much structure to add.&lt;/p&gt;

&lt;p&gt;PromptLedger could have gone further.&lt;/p&gt;

&lt;p&gt;It could have introduced nested collections, custom role registries, marker categories, richer release channels, or project configuration files.&lt;/p&gt;

&lt;p&gt;I avoided that for now.&lt;/p&gt;

&lt;p&gt;The project works best when the concepts stay small:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;ids identify prompt histories&lt;/li&gt;
&lt;li&gt;versions preserve changes over time&lt;/li&gt;
&lt;li&gt;labels provide stronger named semantics&lt;/li&gt;
&lt;li&gt;markers provide lightweight version signals&lt;/li&gt;
&lt;li&gt;collection groups prompt versions&lt;/li&gt;
&lt;li&gt;role explains what kind of prompt artifact a version is&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is enough structure to make a growing prompt library easier to navigate, without making the tool feel like a platform.&lt;/p&gt;

&lt;p&gt;This is also why markers are not labels with another name.&lt;/p&gt;

&lt;p&gt;Labels are useful when you want something closer to a maintained pointer or status. Markers are useful when you want to annotate a version as notable.&lt;/p&gt;

&lt;p&gt;Both can exist because they answer different workflow questions.&lt;/p&gt;

&lt;p&gt;A label asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What does this version represent in a larger workflow?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A marker asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Is this version worth noticing later?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That difference is small, but it keeps the model clean.&lt;/p&gt;




&lt;h2&gt;
  
  
  What did not change
&lt;/h2&gt;

&lt;p&gt;v0.4 does not change the basic philosophy of PromptLedger.&lt;/p&gt;

&lt;p&gt;It is still a small, local-first prompt version control tool.&lt;/p&gt;

&lt;p&gt;It still does not add:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;hosted registry&lt;/li&gt;
&lt;li&gt;prompt execution APIs&lt;/li&gt;
&lt;li&gt;agent tooling&lt;/li&gt;
&lt;li&gt;telemetry pipelines&lt;/li&gt;
&lt;li&gt;cloud sync&lt;/li&gt;
&lt;li&gt;automatic scoring&lt;/li&gt;
&lt;li&gt;evaluation harnesses&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those are not bad ideas in general. They are just outside the current scope of this project.&lt;/p&gt;

&lt;p&gt;PromptLedger is not trying to run prompts, score prompts, deploy prompts, or orchestrate agents.&lt;/p&gt;

&lt;p&gt;It is trying to make prompt changes easier to store, compare, review, and organize.&lt;/p&gt;

&lt;p&gt;That boundary is useful.&lt;/p&gt;

&lt;p&gt;Small scope is not a missing feature here. It is part of the design.&lt;/p&gt;




&lt;h2&gt;
  
  
  Closing thoughts
&lt;/h2&gt;

&lt;p&gt;PromptLedger v0.4 is a usability and organization release.&lt;/p&gt;

&lt;p&gt;It does not radically change what the tool is. Instead, it makes the existing workflow smoother.&lt;/p&gt;

&lt;p&gt;Quick add reduces friction during iteration. Collection and role make prompt libraries easier to navigate. Markers create a lightweight way to remember important versions. Search, list, show, and the read-only UI now expose more of that structure.&lt;/p&gt;

&lt;p&gt;The result is still intentionally modest.&lt;/p&gt;

&lt;p&gt;A local SQLite database. A CLI. A read-only inspection UI. Deterministic exports and reviewable history where possible.&lt;/p&gt;

&lt;p&gt;But the day-to-day workflow feels better now.&lt;/p&gt;

&lt;p&gt;And for a tool like this, that matters.&lt;/p&gt;

&lt;p&gt;Because prompt version control only becomes useful when it is easy enough to use consistently.&lt;/p&gt;




&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;p&gt;PyPI: &lt;a href="https://pypi.org/project/promptledger/" rel="noopener noreferrer"&gt;PyPI&lt;/a&gt;&lt;br&gt;
GitHub: &lt;a href="https://github.com/Ertugrulmutlu/promptledger" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;&lt;br&gt;
LinkedIn: &lt;a href="https://www.linkedin.com/in/ertugrul-mutlu/?locale=en" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt;&lt;br&gt;
Website: &lt;a href="https://ertugrulmutlu.github.io" rel="noopener noreferrer"&gt;Website&lt;/a&gt;&lt;/p&gt;

</description>
      <category>promptengineering</category>
      <category>python</category>
      <category>opensource</category>
      <category>showdev</category>
    </item>
    <item>
      <title>PromptLedger v0.3 — Turning prompt history into a practical review workflow.</title>
      <dc:creator>Ertugrul</dc:creator>
      <pubDate>Sat, 28 Mar 2026 13:37:59 +0000</pubDate>
      <link>https://dev.to/ertugrulmutlu/promptledger-v03-labels-status-and-better-diffs-1ong</link>
      <guid>https://dev.to/ertugrulmutlu/promptledger-v03-labels-status-and-better-diffs-1ong</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Devlog — Part 3&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Turning prompt history into a practical review workflow.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;p&gt;In Part 1, I introduced &lt;strong&gt;PromptLedger&lt;/strong&gt; as a deliberately small, local-first tool for treating prompts like code.&lt;/p&gt;

&lt;p&gt;In Part 2, I added &lt;strong&gt;release semantics&lt;/strong&gt;: labels, label history, and status views that made it easier to answer questions like &lt;em&gt;what is in production right now?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;With &lt;strong&gt;v0.3&lt;/strong&gt;, the next question became harder:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Even if I can diff two prompt versions, can I review them in a way that feels closer to a real release workflow?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That is the focus of this release.&lt;/p&gt;

&lt;p&gt;PromptLedger v0.3 adds a small but practical &lt;strong&gt;Prompt Review&lt;/strong&gt; layer on top of the existing history model — while still staying local-first, SQLite-backed, and intentionally limited in scope.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why a third part?
&lt;/h2&gt;

&lt;p&gt;After the release semantics work in v0.2, the project could already answer questions like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which prompt does &lt;code&gt;prod&lt;/code&gt; currently point to?&lt;/li&gt;
&lt;li&gt;When was that label changed?&lt;/li&gt;
&lt;li&gt;How does &lt;code&gt;prod&lt;/code&gt; differ from &lt;code&gt;staging&lt;/code&gt;?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But another gap became obvious.&lt;/p&gt;

&lt;p&gt;A raw diff is useful, but in practice people often want a slightly higher-level review:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Did the prompt become stricter?&lt;/li&gt;
&lt;li&gt;Did the tone change?&lt;/li&gt;
&lt;li&gt;Was the output format changed from bullets to JSON?&lt;/li&gt;
&lt;li&gt;Did safety or refusal wording get stronger or weaker?&lt;/li&gt;
&lt;li&gt;Is this a release change or a likely regression risk?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those are not execution questions. They are not observability questions either.&lt;/p&gt;

&lt;p&gt;They are &lt;strong&gt;review questions&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;So instead of adding prompt execution, external APIs, or any hosted layer, I kept the project focused and added a &lt;strong&gt;review workflow built entirely on top of the existing local data&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  The main addition: &lt;code&gt;review&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;The new command is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;promptledger review &lt;span class="nt"&gt;--id&lt;/span&gt; onboarding &lt;span class="nt"&gt;--from&lt;/span&gt; prod &lt;span class="nt"&gt;--to&lt;/span&gt; staging
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This compares two refs — versions or labels — and produces a structured review output that includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;resolved refs and versions&lt;/li&gt;
&lt;li&gt;a semantic summary&lt;/li&gt;
&lt;li&gt;metadata changes&lt;/li&gt;
&lt;li&gt;label context&lt;/li&gt;
&lt;li&gt;warning flags&lt;/li&gt;
&lt;li&gt;a few conservative notes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is deliberately not an evaluation system. It does not score prompts. It does not call a model. It does not guess too much.&lt;/p&gt;

&lt;p&gt;It simply makes a prompt diff easier to interpret.&lt;/p&gt;




&lt;h2&gt;
  
  
  From line diff to semantic summary
&lt;/h2&gt;

&lt;p&gt;Traditional diffs are still useful, and PromptLedger keeps all previous diff modes.&lt;/p&gt;

&lt;p&gt;But v0.3 adds a new summary-oriented mode:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;promptledger diff &lt;span class="nt"&gt;--id&lt;/span&gt; onboarding &lt;span class="nt"&gt;--from&lt;/span&gt; 7 &lt;span class="nt"&gt;--to&lt;/span&gt; 9 &lt;span class="nt"&gt;--mode&lt;/span&gt; summary
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This produces a &lt;strong&gt;heuristic, rule-based semantic summary&lt;/strong&gt; instead of a raw line diff.&lt;/p&gt;

&lt;p&gt;The important design decision here is that the summary is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;local&lt;/li&gt;
&lt;li&gt;deterministic&lt;/li&gt;
&lt;li&gt;transparent&lt;/li&gt;
&lt;li&gt;intentionally conservative&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In other words: it only says something when the change looks clear enough.&lt;/p&gt;

&lt;p&gt;Current summary categories include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;tone changes&lt;/li&gt;
&lt;li&gt;tighter or looser constraints&lt;/li&gt;
&lt;li&gt;output format changes&lt;/li&gt;
&lt;li&gt;broader vs more specific prompts&lt;/li&gt;
&lt;li&gt;safety wording changes&lt;/li&gt;
&lt;li&gt;length requirement changes&lt;/li&gt;
&lt;li&gt;refusal or policy wording changes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is not meant to replace reading the actual prompt.&lt;br&gt;
It is meant to make review faster and more structured.&lt;/p&gt;


&lt;h2&gt;
  
  
  Why heuristics instead of an LLM?
&lt;/h2&gt;

&lt;p&gt;Because using an external model for review would push the project in exactly the wrong direction.&lt;/p&gt;

&lt;p&gt;It would introduce:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;network dependence&lt;/li&gt;
&lt;li&gt;nondeterministic behavior&lt;/li&gt;
&lt;li&gt;more configuration&lt;/li&gt;
&lt;li&gt;harder testing&lt;/li&gt;
&lt;li&gt;less trust in the output&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;PromptLedger is supposed to be inspectable.&lt;br&gt;
If it says &lt;em&gt;“constraints tightened”&lt;/em&gt;, that should come from understandable rules, not hidden inference.&lt;/p&gt;

&lt;p&gt;That made a heuristic system the better fit.&lt;/p&gt;

&lt;p&gt;It is not as flexible as an LLM-based reviewer, but it is much easier to reason about — and much more aligned with the philosophy of the project.&lt;/p&gt;


&lt;h2&gt;
  
  
  Reviews now export cleanly to markdown
&lt;/h2&gt;

&lt;p&gt;Another practical gap in earlier versions was sharing review output.&lt;/p&gt;

&lt;p&gt;Reading a diff in the terminal is fine.&lt;br&gt;
Sharing it in a PR, issue, or internal document is another matter.&lt;/p&gt;

&lt;p&gt;So v0.3 adds markdown export for reviews:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;promptledger &lt;span class="nb"&gt;export &lt;/span&gt;review &lt;span class="nt"&gt;--id&lt;/span&gt; onboarding &lt;span class="nt"&gt;--from&lt;/span&gt; prod &lt;span class="nt"&gt;--to&lt;/span&gt; staging &lt;span class="nt"&gt;--format&lt;/span&gt; md &lt;span class="nt"&gt;--out&lt;/span&gt; review.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The exported markdown is deterministic and structured.&lt;br&gt;
It includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a title&lt;/li&gt;
&lt;li&gt;compared refs&lt;/li&gt;
&lt;li&gt;semantic summary&lt;/li&gt;
&lt;li&gt;text diff note&lt;/li&gt;
&lt;li&gt;metadata changes&lt;/li&gt;
&lt;li&gt;warnings&lt;/li&gt;
&lt;li&gt;label information&lt;/li&gt;
&lt;li&gt;a reviewer notes placeholder&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That makes PromptLedger more useful in real workflows without adding any collaboration backend.&lt;/p&gt;

&lt;p&gt;The file is still just a file.&lt;br&gt;
You can paste it into GitHub, attach it to docs, or keep it locally.&lt;/p&gt;




&lt;h2&gt;
  
  
  Metadata changes are now first-class in reviews
&lt;/h2&gt;

&lt;p&gt;Prompt text is only part of the story.&lt;/p&gt;

&lt;p&gt;A release change may also involve metadata updates:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;reason&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;author&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;tags&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;env&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;metrics&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Earlier versions could already diff metadata, but v0.3 makes metadata changes part of the review object itself.&lt;/p&gt;

&lt;p&gt;That matters because some changes are &lt;strong&gt;metadata-only&lt;/strong&gt;.&lt;br&gt;
In those cases, PromptLedger can now say that clearly instead of pretending there was meaningful prompt drift.&lt;/p&gt;

&lt;p&gt;This is a small feature, but an important one.&lt;br&gt;
It avoids overclaiming, which is one of the easiest ways to make a review tool feel unreliable.&lt;/p&gt;




&lt;h2&gt;
  
  
  Warning flags and likely drift hotspots
&lt;/h2&gt;

&lt;p&gt;Prompt review is not just about summarizing what changed.&lt;br&gt;
It is also about drawing attention to changes that deserve extra care.&lt;/p&gt;

&lt;p&gt;v0.3 adds simple warning flags for cases such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;comparing the same version to itself&lt;/li&gt;
&lt;li&gt;environment changes&lt;/li&gt;
&lt;li&gt;metadata-only changes&lt;/li&gt;
&lt;li&gt;policy or refusal wording changes that may affect behavior drift&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These warnings are not meant to be dramatic.&lt;br&gt;
They are meant to make the review output more useful in practice.&lt;/p&gt;

&lt;p&gt;For example, a wording change around refusal or safety does not automatically mean the prompt got worse — but it probably means a reviewer should read it more carefully.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Python API now returns structured review objects
&lt;/h2&gt;

&lt;p&gt;The review workflow is not just a CLI feature.&lt;/p&gt;

&lt;p&gt;The Python API now exposes review results as structured domain objects rather than just formatted strings.&lt;/p&gt;

&lt;p&gt;That means callers can programmatically access:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;resolved refs&lt;/li&gt;
&lt;li&gt;semantic summary items&lt;/li&gt;
&lt;li&gt;metadata changes&lt;/li&gt;
&lt;li&gt;warnings&lt;/li&gt;
&lt;li&gt;notes&lt;/li&gt;
&lt;li&gt;label context&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This keeps the CLI and the API aligned while also making formatting a separate concern.&lt;/p&gt;

&lt;p&gt;That separation turned out to be one of the cleaner changes in this version:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;review logic lives in one place&lt;/li&gt;
&lt;li&gt;rendering logic lives elsewhere&lt;/li&gt;
&lt;li&gt;markdown export and terminal rendering are both built on the same review result&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Small project, but still worth keeping modular.&lt;/p&gt;




&lt;h2&gt;
  
  
  UI update: review without write access
&lt;/h2&gt;

&lt;p&gt;The Streamlit UI is still read-only.&lt;br&gt;
That did not change.&lt;/p&gt;

&lt;p&gt;What changed is that the comparison view now surfaces review information more clearly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;semantic summary&lt;/li&gt;
&lt;li&gt;warnings&lt;/li&gt;
&lt;li&gt;metadata diff&lt;/li&gt;
&lt;li&gt;side-by-side prompt comparison&lt;/li&gt;
&lt;li&gt;line diff&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This keeps the UI aligned with the CLI review flow without turning it into an editor.&lt;/p&gt;

&lt;p&gt;That constraint still matters.&lt;br&gt;
The UI is there to inspect history, not to mutate it.&lt;/p&gt;




&lt;h2&gt;
  
  
  What did &lt;em&gt;not&lt;/em&gt; change
&lt;/h2&gt;

&lt;p&gt;Just as important as the new features is what was left out.&lt;/p&gt;

&lt;p&gt;v0.3 does &lt;strong&gt;not&lt;/strong&gt; add:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a hosted registry&lt;/li&gt;
&lt;li&gt;prompt execution APIs&lt;/li&gt;
&lt;li&gt;agent tooling&lt;/li&gt;
&lt;li&gt;telemetry pipelines&lt;/li&gt;
&lt;li&gt;tracing dashboards&lt;/li&gt;
&lt;li&gt;cloud sync&lt;/li&gt;
&lt;li&gt;automatic scoring&lt;/li&gt;
&lt;li&gt;evaluation harnesses&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There are already plenty of tools going in those directions.&lt;/p&gt;

&lt;p&gt;PromptLedger is still trying to do one narrower thing well:&lt;br&gt;
&lt;strong&gt;store, compare, review, and export prompt changes locally.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  No schema expansion was needed
&lt;/h2&gt;

&lt;p&gt;One part of this release that I particularly liked: the review workflow did not require turning the database into something more complicated.&lt;/p&gt;

&lt;p&gt;SQLite remains the single source of truth.&lt;br&gt;
The review layer is generated from existing prompt versions, labels, and metadata.&lt;/p&gt;

&lt;p&gt;That kept the implementation smaller and the migration story simpler.&lt;/p&gt;

&lt;p&gt;Not every useful feature needs a bigger schema.&lt;br&gt;
Sometimes the better move is to extract more value from the structure that is already there.&lt;/p&gt;




&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;v0.3 did not try to make PromptLedger smarter in a flashy way.&lt;br&gt;
It tried to make it &lt;strong&gt;more reviewable&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The result is still a local tool.&lt;br&gt;
Still inspectable.&lt;br&gt;
Still deterministic where possible.&lt;br&gt;
Still intentionally limited.&lt;/p&gt;

&lt;p&gt;But now it is easier to answer a more realistic question:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Not just “what changed?” — but “how should I review this change before I move it forward?”&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That is a better place for the project to be.&lt;/p&gt;




&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;p&gt;PyPI: &lt;a href="https://pypi.org/project/promptledger/" rel="noopener noreferrer"&gt;PyPI&lt;/a&gt;&lt;br&gt;
GitHub: &lt;a href="https://github.com/Ertugrulmutlu/promptledger" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;&lt;br&gt;
LinkedIn: &lt;a href="https://www.linkedin.com/in/ertugrul-mutlu/?locale=en" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt;&lt;br&gt;
Website: &lt;a href="https://ertugrulmutlu.github.io" rel="noopener noreferrer"&gt;Website&lt;/a&gt;&lt;/p&gt;

</description>
      <category>promptengineering</category>
      <category>python</category>
      <category>opensource</category>
      <category>showdev</category>
    </item>
    <item>
      <title>DataLens: A Read-Only Image Dataset Sanity Checker</title>
      <dc:creator>Ertugrul</dc:creator>
      <pubDate>Tue, 20 Jan 2026 13:43:50 +0000</pubDate>
      <link>https://dev.to/ertugrulmutlu/datalens-a-read-only-image-dataset-sanity-checker-2do0</link>
      <guid>https://dev.to/ertugrulmutlu/datalens-a-read-only-image-dataset-sanity-checker-2do0</guid>
      <description>&lt;h1&gt;
  
  
  DataLens: A Read‑Only Image Dataset Sanity Checker
&lt;/h1&gt;

&lt;p&gt;Training a model rarely fails loudly.&lt;/p&gt;

&lt;p&gt;Most of the time, it &lt;em&gt;kind of works&lt;/em&gt; — loss decreases, accuracy moves, but the results feel unstable, brittle, or just wrong.&lt;/p&gt;

&lt;p&gt;In my experience, when that happens, the root cause is often not the model, but the &lt;strong&gt;dataset&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;So I built &lt;strong&gt;DataLens&lt;/strong&gt;: a lightweight, read‑only sanity checker for image datasets.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Problem: Silent Dataset Failures
&lt;/h2&gt;

&lt;p&gt;Before training even starts, datasets often contain issues like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Corrupted or unreadable images&lt;/li&gt;
&lt;li&gt;Duplicate or near‑duplicate samples&lt;/li&gt;
&lt;li&gt;Broken CSV → image mappings&lt;/li&gt;
&lt;li&gt;Large numbers of orphan images&lt;/li&gt;
&lt;li&gt;Severe class imbalance&lt;/li&gt;
&lt;li&gt;Extremely small images or extreme aspect ratios&lt;/li&gt;
&lt;li&gt;Mixed image modes (RGB / RGBA / grayscale)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of these necessarily crash training.&lt;br&gt;
They just quietly degrade everything downstream.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why Read‑Only Matters
&lt;/h2&gt;

&lt;p&gt;Many tools try to &lt;strong&gt;auto‑fix&lt;/strong&gt; datasets.&lt;/p&gt;

&lt;p&gt;I deliberately didn’t.&lt;/p&gt;

&lt;p&gt;DataLens follows a simple rule:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Inspect, don’t mutate.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;ul&gt;
&lt;li&gt;No files are moved&lt;/li&gt;
&lt;li&gt;No labels are rewritten&lt;/li&gt;
&lt;li&gt;No assumptions are silently applied&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The tool’s job is to surface problems clearly, so &lt;em&gt;you&lt;/em&gt; can decide what to do next.&lt;/p&gt;




&lt;h2&gt;
  
  
  What DataLens Does
&lt;/h2&gt;

&lt;p&gt;DataLens is a Streamlit‑based audit tool with two modes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Mode A — Images Only
&lt;/h3&gt;

&lt;p&gt;For raw image folders:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Recursively scans images&lt;/li&gt;
&lt;li&gt;Detects corrupted files&lt;/li&gt;
&lt;li&gt;Finds exact and near‑duplicate images&lt;/li&gt;
&lt;li&gt;Optionally infers classes from subfolders&lt;/li&gt;
&lt;li&gt;Correctly handles &lt;strong&gt;unlabeled datasets&lt;/strong&gt; (no fake “missing label” warnings)&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Mode B — Images + Labels (CSV)
&lt;/h3&gt;

&lt;p&gt;For supervised datasets:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Robust CSV reading (UTF‑8 with fallback)&lt;/li&gt;
&lt;li&gt;Automatic filename &amp;amp; label column detection&lt;/li&gt;
&lt;li&gt;Support for IDs without file extensions&lt;/li&gt;
&lt;li&gt;Optional label normalization&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Coverage analysis:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;em&gt;How many CSV rows actually resolve to images?&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;


&lt;/li&gt;

&lt;li&gt;

&lt;p&gt;Orphan analysis:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;em&gt;How many images are never referenced by the CSV?&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;


&lt;/li&gt;

&lt;/ul&gt;




&lt;h2&gt;
  
  
  Duplicate Detection That Actually Helps
&lt;/h2&gt;

&lt;p&gt;Exact duplicates are easy.&lt;br&gt;
Near‑duplicates are not.&lt;/p&gt;

&lt;p&gt;DataLens supports three methods:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;sha256&lt;/strong&gt; — byte‑exact duplicates&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;quick hash&lt;/strong&gt; — fast approximation&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;pHash&lt;/strong&gt; — visually similar images (resized, recompressed, slightly cropped)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is especially useful for datasets collected via scraping or merging multiple sources.&lt;/p&gt;




&lt;h2&gt;
  
  
  Data Hygiene Warnings
&lt;/h2&gt;

&lt;p&gt;Beyond basic checks, DataLens flags issues that usually show up &lt;em&gt;too late&lt;/em&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Very small images (e.g. &amp;lt;64px)&lt;/li&gt;
&lt;li&gt;Extreme aspect ratios&lt;/li&gt;
&lt;li&gt;High RGBA share (alpha channel surprises)&lt;/li&gt;
&lt;li&gt;High image mode variance&lt;/li&gt;
&lt;li&gt;Extension mismatches between CSV references and actual files&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are the kinds of things that quietly break training pipelines or augmentations.&lt;/p&gt;




&lt;h2&gt;
  
  
  Outputs You Can Share
&lt;/h2&gt;

&lt;p&gt;After a run, DataLens produces:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;An &lt;strong&gt;interactive dashboard&lt;/strong&gt; (Streamlit)&lt;/li&gt;
&lt;li&gt;A deterministic &lt;code&gt;dataset_report.md&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;An &lt;code&gt;issues.csv&lt;/code&gt; containing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;missing images&lt;/li&gt;
&lt;li&gt;orphan images&lt;/li&gt;
&lt;li&gt;corrupted files&lt;/li&gt;
&lt;li&gt;duplicate groups&lt;/li&gt;
&lt;/ul&gt;


&lt;/li&gt;

&lt;/ul&gt;

&lt;p&gt;The report is designed to be:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;commit‑friendly&lt;/li&gt;
&lt;li&gt;reviewable&lt;/li&gt;
&lt;li&gt;attachable to issues or PRs&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Design Principles
&lt;/h2&gt;

&lt;p&gt;I kept the scope intentionally tight:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Read‑only&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Deterministic&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Transparent&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Pre‑training focused&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is not a dataset cleaning tool.&lt;br&gt;
It’s a &lt;strong&gt;dataset inspection tool&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  When I Use DataLens
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Before starting any new training run&lt;/li&gt;
&lt;li&gt;When receiving datasets from external sources&lt;/li&gt;
&lt;li&gt;When debugging unstable or suspicious training behavior&lt;/li&gt;
&lt;li&gt;As a lightweight QA step before investing GPU hours&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Final Thought
&lt;/h2&gt;

&lt;p&gt;We spend a lot of time tuning models.&lt;/p&gt;

&lt;p&gt;But models are only as good as the data we feed them — and bad data usually doesn’t announce itself.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;DataLens is my way of making datasets talk.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/Ertugrulmutlu/DataLens" rel="noopener noreferrer"&gt;https://github.com/Ertugrulmutlu/DataLens&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LinkedIn:&lt;/strong&gt; &lt;a href="https://www.linkedin.com/in/ertugrul-mutlu/" rel="noopener noreferrer"&gt;https://www.linkedin.com/in/ertugrul-mutlu/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Personal Website:&lt;/strong&gt; &lt;a href="https://ertugrulmutlu.github.io" rel="noopener noreferrer"&gt;https://ertugrulmutlu.github.io&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>machinelearning</category>
      <category>showdev</category>
      <category>testing</category>
      <category>tooling</category>
    </item>
    <item>
      <title>PromptLedger v0.2 — Labels, Status, and Better Diffs</title>
      <dc:creator>Ertugrul</dc:creator>
      <pubDate>Tue, 13 Jan 2026 14:44:54 +0000</pubDate>
      <link>https://dev.to/ertugrulmutlu/promptledger-v02-labels-status-and-better-diffs-137j</link>
      <guid>https://dev.to/ertugrulmutlu/promptledger-v02-labels-status-and-better-diffs-137j</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Devlog — Part 2&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;What changed after the first release, and why those changes matter in practice.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;p&gt;In the first part, I introduced &lt;strong&gt;PromptLedger&lt;/strong&gt; as a deliberately small, local-first tool for treating prompts like code.&lt;/p&gt;

&lt;p&gt;Since then, the project has moved from a minimal prompt history tracker to something closer to a &lt;strong&gt;prompt release ledger&lt;/strong&gt; — without adding servers, agents, or execution layers.&lt;/p&gt;

&lt;p&gt;This post covers what changed in &lt;strong&gt;v0.2&lt;/strong&gt;, why those changes exist, and how they affect real prompt workflows.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why a second part?
&lt;/h2&gt;

&lt;p&gt;After publishing Part 1, the most common follow-up questions were:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;em&gt;Which prompt is actually in production right now?&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;em&gt;How do I compare prod vs staging without remembering version numbers?&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Can I see when a release pointer changed?&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Answering these questions required &lt;strong&gt;release semantics&lt;/strong&gt;, not more prompt editing features.&lt;/p&gt;

&lt;p&gt;That is the focus of v0.2.&lt;/p&gt;




&lt;h2&gt;
  
  
  Label history: prompts need an audit trail
&lt;/h2&gt;

&lt;p&gt;In v0.1, labels existed as movable pointers, similar to git tags. However, once a label moved, its previous state was lost.&lt;/p&gt;

&lt;p&gt;In v0.2, &lt;strong&gt;every label update is recorded in an append-only history log&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;This means you can now answer questions like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What prompt was in production yesterday?&lt;/li&gt;
&lt;li&gt;When did &lt;code&gt;prod&lt;/code&gt; move from version 7 to version 9?
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;promptledger label &lt;span class="nb"&gt;history&lt;/span&gt; &lt;span class="nt"&gt;--id&lt;/span&gt; onboarding
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Under the hood, label history is implemented as a separate &lt;code&gt;label_events&lt;/code&gt; table. Label pointers remain mutable, but the history is immutable.&lt;/p&gt;

&lt;p&gt;This keeps the system simple while adding real auditability.&lt;/p&gt;




&lt;h2&gt;
  
  
  Label-based diff: stop thinking in version numbers
&lt;/h2&gt;

&lt;p&gt;Version numbers are great for storage, but humans think in environments.&lt;/p&gt;

&lt;p&gt;v0.2 allows diffs to be expressed in terms of &lt;strong&gt;labels&lt;/strong&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;promptledger diff &lt;span class="nt"&gt;--id&lt;/span&gt; onboarding &lt;span class="nt"&gt;--from&lt;/span&gt; prod &lt;span class="nt"&gt;--to&lt;/span&gt; staging
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This resolves labels to their underlying versions and performs a normal diff.&lt;/p&gt;

&lt;p&gt;The important detail is that &lt;strong&gt;nothing new is stored&lt;/strong&gt;. Labels are only references; diffs always operate on immutable prompt versions.&lt;/p&gt;




&lt;h2&gt;
  
  
  Status command: a one-line overview
&lt;/h2&gt;

&lt;p&gt;Another recurring problem was simply understanding the current state of prompts.&lt;/p&gt;

&lt;p&gt;The new &lt;code&gt;status&lt;/code&gt; command provides a concise, git-style overview:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;promptledger status
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For each prompt, it shows:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The latest version number&lt;/li&gt;
&lt;li&gt;The timestamp of that version&lt;/li&gt;
&lt;li&gt;Active labels and the versions they point to&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is intentionally read-only and summary-focused. If you need details, you still drill down using &lt;code&gt;list&lt;/code&gt;, &lt;code&gt;show&lt;/code&gt;, or &lt;code&gt;diff&lt;/code&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  Better diff modes
&lt;/h2&gt;

&lt;p&gt;Prompt changes are not always best reviewed line-by-line.&lt;/p&gt;

&lt;p&gt;v0.2 introduces multiple diff modes built on top of Python’s &lt;code&gt;difflib&lt;/code&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;unified&lt;/code&gt; — the default, git-style view&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;context&lt;/code&gt; — useful for wider structural changes&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;ndiff&lt;/code&gt; — character-level insight for small edits&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;metadata&lt;/code&gt; — diff &lt;strong&gt;only&lt;/strong&gt; metadata (&lt;code&gt;reason&lt;/code&gt;, &lt;code&gt;tags&lt;/code&gt;, &lt;code&gt;env&lt;/code&gt;, &lt;code&gt;metrics&lt;/code&gt;)
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;promptledger diff &lt;span class="nt"&gt;--id&lt;/span&gt; onboarding &lt;span class="nt"&gt;--from&lt;/span&gt; 1 &lt;span class="nt"&gt;--to&lt;/span&gt; 2 &lt;span class="nt"&gt;--mode&lt;/span&gt; metadata
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Metadata diffs are always rendered in unified format to keep them readable and deterministic.&lt;/p&gt;




&lt;h2&gt;
  
  
  UI update: history without write access
&lt;/h2&gt;

&lt;p&gt;The Streamlit UI remains strictly read-only, but v0.2 adds &lt;strong&gt;label history visibility&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;You can now:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Inspect prompt timelines&lt;/li&gt;
&lt;li&gt;See which labels point where&lt;/li&gt;
&lt;li&gt;Review label movements over time&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All mutations still go through the CLI or Python API.&lt;/p&gt;

&lt;p&gt;This constraint is intentional: the UI is for inspection, not experimentation.&lt;/p&gt;




&lt;h2&gt;
  
  
  What did &lt;em&gt;not&lt;/em&gt; change
&lt;/h2&gt;

&lt;p&gt;Several things were deliberately left out:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;No prompt execution or playground&lt;/li&gt;
&lt;li&gt;No agent framework&lt;/li&gt;
&lt;li&gt;No cloud sync or backend service&lt;/li&gt;
&lt;li&gt;No automatic evaluation or scoring&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;PromptLedger is still a &lt;strong&gt;ledger&lt;/strong&gt;, not an environment.&lt;/p&gt;




&lt;h2&gt;
  
  
  Migration and compatibility
&lt;/h2&gt;

&lt;p&gt;v0.2 introduces a database schema migration, but it is fully backwards compatible.&lt;/p&gt;

&lt;p&gt;Existing databases are upgraded in place, with no data loss.&lt;/p&gt;

&lt;p&gt;The new tables add information, but do not change existing semantics.&lt;/p&gt;




&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;v0.2 did not make PromptLedger bigger — it made it clearer.&lt;/p&gt;

&lt;p&gt;By adding release semantics, auditability, and better inspection tools, prompts can now be reviewed and promoted with the same discipline as code.&lt;/p&gt;

&lt;p&gt;No servers. No dashboards. No magic.&lt;/p&gt;

&lt;p&gt;Just history, diffs, and intent — stored locally.&lt;/p&gt;




&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;p&gt;Pypi : &lt;a href="https://pypi.org/project/promptledger/" rel="noopener noreferrer"&gt;Pypi&lt;/a&gt;&lt;br&gt;
Github : &lt;a href="https://github.com/Ertugrulmutlu/promptledger" rel="noopener noreferrer"&gt;Github&lt;/a&gt;&lt;br&gt;
My Linkedin : &lt;a href="https://www.linkedin.com/in/ertu%C4%9Frul-mutlu/" rel="noopener noreferrer"&gt;Linkedin&lt;/a&gt;&lt;br&gt;
My Website : &lt;a href="https://ertugrulmutlu.github.io" rel="noopener noreferrer"&gt;Website&lt;/a&gt;&lt;/p&gt;

</description>
      <category>promptengineering</category>
      <category>python</category>
      <category>opensource</category>
      <category>showdev</category>
    </item>
    <item>
      <title>Hysteresis in Neural Networks — Part 1</title>
      <dc:creator>Ertugrul</dc:creator>
      <pubDate>Fri, 09 Jan 2026 14:09:14 +0000</pubDate>
      <link>https://dev.to/ertugrulmutlu/hysteresis-in-neural-networks-part-1-3c03</link>
      <guid>https://dev.to/ertugrulmutlu/hysteresis-in-neural-networks-part-1-3c03</guid>
      <description>&lt;h2&gt;
  
  
  Training Order Is Not Innocent
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;What if seeing the same data is not enough?&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A common, almost implicit assumption in machine learning is that &lt;strong&gt;if a model is trained on the same dataset with the same architecture and optimizer, it should end up learning essentially the same thing&lt;/strong&gt;. Training order is usually treated as an implementation detail — a convenience, not a defining factor.&lt;/p&gt;

&lt;p&gt;In this post, I show a simple but striking counterexample:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Even when two neural networks see exactly the same data, the &lt;em&gt;order&lt;/em&gt; in which the data is presented can determine what the model permanently remembers and what it completely forgets.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is the first part of a short series on &lt;em&gt;hysteresis in neural networks&lt;/em&gt;. Here, I focus only on observable behavior (accuracy and forgetting). In the next part, we will look inside the model and explain &lt;em&gt;why&lt;/em&gt; this happens geometrically.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Core Question
&lt;/h2&gt;

&lt;p&gt;Assume we have two datasets, &lt;strong&gt;A&lt;/strong&gt; and &lt;strong&gt;B&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;We train the same model in two different ways:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;SAB&lt;/strong&gt;: first on A, then on B&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SBA&lt;/strong&gt;: first on B, then on A&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Crucially:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The architecture is identical&lt;/li&gt;
&lt;li&gt;The optimizer and hyperparameters are identical&lt;/li&gt;
&lt;li&gt;The random seed is fixed&lt;/li&gt;
&lt;li&gt;The union of the data is the same: A ∪ B&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The only difference is &lt;strong&gt;chronological order&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The standard intuition is that this should not matter.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This intuition is wrong.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Experimental Setup (Intentionally Simple)
&lt;/h2&gt;

&lt;p&gt;To avoid hiding effects behind complexity, I used a deliberately minimal setup.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Dataset&lt;/strong&gt;: MNIST&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Split&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A = digits {0,1,2,3,4}&lt;/li&gt;
&lt;li&gt;B = digits {5,6,7,8,9}&lt;/li&gt;
&lt;/ul&gt;


&lt;/li&gt;

&lt;li&gt;&lt;p&gt;&lt;strong&gt;Model&lt;/strong&gt;: small CNN + MLP head&lt;/p&gt;&lt;/li&gt;

&lt;li&gt;&lt;p&gt;&lt;strong&gt;Training&lt;/strong&gt;: 20 epochs total (10 per phase)&lt;/p&gt;&lt;/li&gt;

&lt;/ul&gt;

&lt;p&gt;Several things were &lt;em&gt;explicitly disabled&lt;/em&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;No Batch Normalization&lt;/li&gt;
&lt;li&gt;No data augmentation&lt;/li&gt;
&lt;li&gt;Deterministic initialization (fixed seed)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This ensures that any difference we observe is not an artifact of randomness or regularization, but a consequence of training order alone.&lt;/p&gt;




&lt;h2&gt;
  
  
  What Happens During Training?
&lt;/h2&gt;

&lt;p&gt;Let’s start with the &lt;strong&gt;SAB&lt;/strong&gt; scenario: the model learns A first, then B.&lt;/p&gt;

&lt;h3&gt;
  
  
  SAB: A → B
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Observed behavior&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;During the first phase, accuracy on A quickly rises to ~99%&lt;/li&gt;
&lt;li&gt;Accuracy on B remains near 0% (expected)&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;After switching to B:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Accuracy on B rises to ~99%&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Accuracy on A collapses to 0%&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;


&lt;/li&gt;

&lt;/ul&gt;

&lt;h3&gt;
  
  
  SBA: B → A
&lt;/h3&gt;

&lt;p&gt;The mirror experiment produces the mirror result:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Observed behavior&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;During the first phase, accuracy on B rises to ~99%&lt;/li&gt;
&lt;li&gt;Accuracy on A remains near 0%&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;After switching to A:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Accuracy on A rises to ~99%&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Accuracy on B collapses to 0%&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;


&lt;/li&gt;

&lt;/ul&gt;




&lt;h2&gt;
  
  
  Accuracy Curves
&lt;/h2&gt;

&lt;h3&gt;
  
  
  SAB accuracy over epochs
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Frp6w9v27qbcusko1htpz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Frp6w9v27qbcusko1htpz.png" alt="SAB accuracy over epochs" width="800" height="600"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Description:&lt;/em&gt; Line plot showing &lt;code&gt;acc_A&lt;/code&gt;, &lt;code&gt;acc_B&lt;/code&gt;, and &lt;code&gt;acc_full&lt;/code&gt; across epochs. A vertical dashed line marks the phase transition (A → B). Accuracy on A drops sharply after the transition.&lt;/p&gt;




&lt;h3&gt;
  
  
  SBA accuracy over epochs
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Ff82nk9au3a6u99uwt4mf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Ff82nk9au3a6u99uwt4mf.png" alt="SBA accuracy over epochs" width="800" height="600"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Description:&lt;/em&gt; Same plot as above, but for SBA. Accuracy on B drops sharply after switching to A.&lt;/p&gt;




&lt;h2&gt;
  
  
  A Subtle but Important Observation
&lt;/h2&gt;

&lt;p&gt;Despite the dramatic forgetting, &lt;strong&gt;overall accuracy on the full test set remains around ~50% in both cases&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;This is not a contradiction.&lt;/p&gt;

&lt;p&gt;Each model performs extremely well on &lt;em&gt;half&lt;/em&gt; of the classes and completely fails on the other half. Aggregated metrics hide this asymmetry.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Two models can have similar overall accuracy while representing fundamentally different worlds internally.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Quantifying the Effect: Hysteresis Loss
&lt;/h2&gt;

&lt;p&gt;We can define a simple, order-dependent metric:&lt;/p&gt;

&lt;p&gt;[\mathcal{L}_{\text{hyst}}(A) = |\text{Acc}_A(SAB) - \text{Acc}_A(SBA)|]&lt;/p&gt;

&lt;p&gt;In this experiment:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Acc_A(SAB) ≈ 0.00&lt;/li&gt;
&lt;li&gt;Acc_A(SBA) ≈ 0.99&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This yields a hysteresis loss close to &lt;strong&gt;1.0&lt;/strong&gt;, i.e. the maximum possible difference.&lt;/p&gt;

&lt;p&gt;The same holds symmetrically for dataset B.&lt;/p&gt;




&lt;h3&gt;
  
  
  Hysteresis summary
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fniatonl4y3kwpf2iq8b5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fniatonl4y3kwpf2iq8b5.png" alt="Hysteresis summary" width="800" height="600"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Description:&lt;/em&gt; Bar chart showing absolute accuracy differences between SAB and SBA. &lt;br&gt;
Hysteresis for the individual subsets A and B is near-maximal, while hysteresis in the aggregate (full test accuracy) is close to zero and visually compressed due to the shared scale.&lt;/p&gt;

&lt;p&gt;This highlights a key point: global performance metrics can remain almost invariant, even when class-conditional behavior is maximally path-dependent.&lt;/p&gt;




&lt;h2&gt;
  
  
  What This Means
&lt;/h2&gt;

&lt;p&gt;This result shows that training order is not just an optimization detail.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The network does not converge to a single, order-independent solution&lt;/li&gt;
&lt;li&gt;Learning leaves &lt;em&gt;path-dependent traces&lt;/em&gt; in weight space&lt;/li&gt;
&lt;li&gt;Once the model commits to one subset, returning to a balanced representation is not trivial&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This behavior is strongly reminiscent of &lt;strong&gt;hysteresis in physical systems&lt;/strong&gt;, where the final state depends on the path taken, not only on the endpoint.&lt;/p&gt;




&lt;h2&gt;
  
  
  What This Post Does &lt;em&gt;Not&lt;/em&gt; Explain (Yet)
&lt;/h2&gt;

&lt;p&gt;This post only shows &lt;em&gt;that&lt;/em&gt; hysteresis exists.&lt;/p&gt;

&lt;p&gt;It does &lt;strong&gt;not&lt;/strong&gt; explain:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Whether the two final models lie in different basins of attraction&lt;/li&gt;
&lt;li&gt;Whether their internal representations are aligned or incompatible&lt;/li&gt;
&lt;li&gt;Whether one can smoothly interpolate between them without loss spikes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These questions require looking at the &lt;strong&gt;geometry of the weight space&lt;/strong&gt;, not just accuracy curves.&lt;/p&gt;

&lt;p&gt;That is exactly what Part 2 will address.&lt;/p&gt;




&lt;h2&gt;
  
  
  Reproducibility
&lt;/h2&gt;

&lt;p&gt;All experiments were run with a fixed configuration and deterministic setup. The code used for this post (FAZ 1) is available on GitHub:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/Ertugrulmutlu/hysteresis-neural-networks" rel="noopener noreferrer"&gt;Hysteresis in Neural Networks&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The geometric analysis (weight trajectories, representation similarity, interpolation barriers) will be released together with the next post.&lt;/p&gt;




&lt;h2&gt;
  
  
  Closing Thoughts
&lt;/h2&gt;

&lt;p&gt;If training order alone can erase entire subsets of knowledge, then:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Curriculum learning has hidden costs&lt;/li&gt;
&lt;li&gt;"Same data" does not imply "same model"&lt;/li&gt;
&lt;li&gt;Optimization is not just minimization — it is a &lt;em&gt;history-dependent process&lt;/em&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In the next post, we will open the model and examine how these irreversible choices are encoded in the geometry of neural manifolds.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Part 2: Inside the Weight Space — Geometry of Hysteresis&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;p&gt;Linkedin: &lt;a href="https://www.linkedin.com/in/ertugrul-mutlu/" rel="noopener noreferrer"&gt;Linkedin&lt;/a&gt;&lt;br&gt;
Github: &lt;a href="https://github.com/Ertugrulmutlu" rel="noopener noreferrer"&gt;Github&lt;/a&gt;&lt;br&gt;
Website: &lt;a href="https://ertugrulmutlu.github.io" rel="noopener noreferrer"&gt;Website&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>computerscience</category>
      <category>deeplearning</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
