<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Milivoje Simonović</title>
    <description>The latest articles on DEV Community by Milivoje Simonović (@milivoje_simonovi_ddc92b).</description>
    <link>https://dev.to/milivoje_simonovi_ddc92b</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4156777%2F1f8f7c55-fb47-498b-b7d3-373f918c25ed.png</url>
      <title>DEV Community: Milivoje Simonović</title>
      <link>https://dev.to/milivoje_simonovi_ddc92b</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/milivoje_simonovi_ddc92b"/>
    <language>en</language>
    <item>
      <title>I converted 126 tree models to ONNX by following the docs. Here is what changed.</title>
      <dc:creator>Milivoje Simonović</dc:creator>
      <pubDate>Sat, 03 Oct 2026 19:27:58 +0000</pubDate>
      <link>https://dev.to/milivoje_simonovi_ddc92b/i-converted-126-tree-models-to-onnx-by-following-the-docs-here-is-what-changed-3a3j</link>
      <guid>https://dev.to/milivoje_simonovi_ddc92b/i-converted-126-tree-models-to-onnx-by-following-the-docs-here-is-what-changed-3a3j</guid>
      <description>&lt;p&gt;I wanted to know how often the documented way of converting a tree model to ONNX produces a file that is not the same model. So I built the pairs the way the skl2onnx and onnxmltools documentation shows, with default settings and float32 input, and checked every one with an equivalence checker.&lt;/p&gt;

&lt;p&gt;The setup: 7 datasets (iris, wine, breast cancer, digits, diabetes, a synthetic regression, and a synthetic table of integer counts), 6 model types, 3 sizes each (10 trees at depth 3, 50 at depth 6, 100 at depth 8). That is 126 pairs. Each pair was checked twice, once on finite inputs only and once on all inputs including NaN.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model type&lt;/th&gt;
&lt;th&gt;Finite inputs&lt;/th&gt;
&lt;th&gt;Including NaN&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;XGBoost via onnxmltools&lt;/td&gt;
&lt;td&gt;21 equivalent&lt;/td&gt;
&lt;td&gt;21 equivalent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;scikit-learn GradientBoosting&lt;/td&gt;
&lt;td&gt;21 equivalent&lt;/td&gt;
&lt;td&gt;21 equivalent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;scikit-learn RandomForest&lt;/td&gt;
&lt;td&gt;21 equivalent&lt;/td&gt;
&lt;td&gt;21 not equivalent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;scikit-learn ExtraTrees&lt;/td&gt;
&lt;td&gt;21 equivalent&lt;/td&gt;
&lt;td&gt;21 not equivalent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;scikit-learn DecisionTree&lt;/td&gt;
&lt;td&gt;21 equivalent&lt;/td&gt;
&lt;td&gt;21 not equivalent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LightGBM via onnxmltools&lt;/td&gt;
&lt;td&gt;6 equivalent, 15 not&lt;/td&gt;
&lt;td&gt;6 equivalent, 15 not&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;On real data, none of this shows. I ran every original and every ONNX file on the rows of its own dataset, 19,908 rows in total across the 126 pairs, and not one output differed. A test set check would have passed all 126.&lt;/p&gt;

&lt;h2&gt;
  
  
  LightGBM: one float32 value goes the wrong way
&lt;/h2&gt;

&lt;p&gt;LightGBM stores split thresholds as float64. The ONNX file stores them as float32. When a threshold is a round decimal, the rounding can move it past the one input value that equals that decimal.&lt;/p&gt;

&lt;p&gt;In the breast cancer model with 10 trees, the first split on feature 23 is &lt;code&gt;x &amp;lt;= 868.2000000000002&lt;/code&gt; in LightGBM and &lt;code&gt;x &amp;lt;= 868.2000122070312&lt;/code&gt; in the ONNX file. The input 868.2, as float32, is exactly 868.2000122070312. LightGBM sends it right (it is larger than the threshold), the ONNX file sends it left (it is equal to the threshold).&lt;/p&gt;

&lt;p&gt;I took the input the checker reported and ran it through the real LightGBM and the real onnxruntime. LightGBM predicts class 0 (raw score -0.853). The ONNX file predicts class 1 (probability 0.786). The raw scores differ by 2.156, which is what the report said was the largest possible difference. For 9 of the 15 pairs the report says the predicted class can change.&lt;/p&gt;

&lt;p&gt;The affected band is tiny: one float32 value per threshold. That is why 19,908 real rows never hit it. It is also why it is a real problem when your features are decimals like prices or measurements, because round decimals are exactly the values people type in.&lt;/p&gt;

&lt;p&gt;The 6 LightGBM pairs that were fine are the digits and the integer count datasets, where every threshold is exactly representable.&lt;/p&gt;

&lt;h2&gt;
  
  
  scikit-learn trees and forests: NaN goes a different way
&lt;/h2&gt;

&lt;p&gt;scikit-learn accepts NaN at prediction time for trees and forests, even if the model never saw a NaN in training. According to its documentation, such samples are sent to the child with the most training samples. The converted ONNX file sends NaN a fixed way instead.&lt;/p&gt;

&lt;p&gt;Breast cancer, one decision tree, one NaN in the input: scikit-learn returns probabilities [0.0, 1.0], the ONNX file returns [1.0, 0.0]. The labels are opposite. On the wine random forest, an input with four NaN values gets class 2 from scikit-learn and class 0 from the ONNX file. For 45 of the 63 scikit-learn pairs the report says the predicted class can change.&lt;/p&gt;

&lt;p&gt;This only matters if NaN can reach your model, for example from a failed join or a missing sensor value. If you impute upstream, it does not. But then nothing is checking that for you either.&lt;/p&gt;

&lt;h2&gt;
  
  
  What was fine
&lt;/h2&gt;

&lt;p&gt;All 42 XGBoost and scikit-learn GradientBoosting pairs were equivalent, with and without NaN. I also tried an old XGBoost report (an issue about &lt;code&gt;base_score&lt;/code&gt; in onnxruntime) and it no longer reproduces with current versions. Those converters did the right thing in these settings.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits
&lt;/h2&gt;

&lt;p&gt;These are models I trained on built-in datasets following the documentation. They are not third-party production models, so this says how the documented recipes behave, not how common the problem is in the wild. The checker supports numeric XGBoost, LightGBM and scikit-learn tree models and ONNX TreeEnsemble operators up to opset 3. Categorical splits and HistGradientBoosting are not supported, so they were not part of this.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run it yourself
&lt;/h2&gt;

&lt;p&gt;The script is in the repository. A quick run takes under a minute, the full run about 12 minutes.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;leafparity[all] skl2onnx onnxmltools
python audit.py &lt;span class="nt"&gt;--quick&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Script and README: &lt;a href="https://github.com/Leafparity/leafparity/tree/main/examples/audit" rel="noopener noreferrer"&gt;https://github.com/Leafparity/leafparity/tree/main/examples/audit&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I am the author of the checker, Leafparity (open source, Apache 2.0, &lt;a href="https://leafparity.com" rel="noopener noreferrer"&gt;https://leafparity.com&lt;/a&gt;). If you have a model and its ONNX file where a check like this finds something, or finds nothing, I would like to hear about it.&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>python</category>
      <category>onnx</category>
      <category>testing</category>
    </item>
    <item>
      <title>Your ONNX file is a different model</title>
      <dc:creator>Milivoje Simonović</dc:creator>
      <pubDate>Fri, 02 Oct 2026 08:05:34 +0000</pubDate>
      <link>https://dev.to/milivoje_simonovi_ddc92b/your-onnx-file-is-a-different-model-1l3i</link>
      <guid>https://dev.to/milivoje_simonovi_ddc92b/your-onnx-file-is-a-different-model-1l3i</guid>
      <description>&lt;p&gt;The scikit-learn ONNX converter documentation has a tutorial about what happens when you switch a decision tree from float64 to float32. On its test set, the original model and its ONNX copy differ by at most about 156. I took that same example and asked a different question: what is the largest difference that can exist anywhere in the input space? The answer is 556. If missing values are allowed, it is 721.&lt;/p&gt;

&lt;p&gt;The test set was not wrong. It was just never going to find the places where the two models disagree.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a test set cannot see it
&lt;/h2&gt;

&lt;p&gt;A tree model is a set of boxes. Every input falls into exactly one leaf per tree, and the leaf decides the output. When you convert a model to ONNX, the thresholds that define those boxes can change slightly. A threshold stored as a float64 in the original becomes a float32 in the converted file, so it moves by a tiny amount. Inputs that fall between the old and the new threshold now go down the other branch.&lt;/p&gt;

&lt;p&gt;Those inputs live in a sliver a few billionths wide next to each threshold. A random sample of a few thousand test rows will almost never land there. Production traffic will eventually, because real data has prices, ratios and counts that sit exactly on round numbers, and round numbers are where thresholds like to be.&lt;/p&gt;

&lt;p&gt;When an input does land in a sliver, the output does not change by a rounding error. It changes by a whole leaf value, which is where the 156 and the 556 come from.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three places the differences hide
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Threshold rounding.&lt;/strong&gt; The float32 cast described above. It affects models whose thresholds are float64 in the original and float32 after conversion, such as LightGBM and scikit-learn trees.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Zero treated as missing.&lt;/strong&gt; LightGBM can be trained with &lt;code&gt;zero_as_missing=True&lt;/code&gt;, which means an input of exactly 0.0 is handled as a missing value. When I converted such a model with onnxmltools, the converter ignored that setting. An input of exactly zero then took a different path in every tree, and the raw scores differed by up to about 295.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;NaN routing.&lt;/strong&gt; Since scikit-learn 1.3, trees can route NaN to the left or the right child at each node, learned during training. In the converted model, the NaN went the other way at the affected nodes.&lt;/p&gt;

&lt;p&gt;None of these are exotic. Each one comes from a reasonable converter doing a reasonable thing in a place where the two libraries have different rules.&lt;/p&gt;

&lt;h2&gt;
  
  
  Checking it without sampling
&lt;/h2&gt;

&lt;p&gt;Because both models are finite sets of boxes, there is a way to compare them without guessing. Walk the two models together over the whole input space. At every region, ask whether the original and the converted model reach corresponding leaves. If they do everywhere, the models agree for every possible input, up to floating point rounding of the outputs. If they do not, you can list each region where they disagree, the exact set of inputs affected, and the largest difference.&lt;/p&gt;

&lt;p&gt;For every problem found, you also want a concrete input you can run through both real runtimes, so nobody has to trust the analysis. If the XGBoost, LightGBM or scikit-learn prediction and the onnxruntime prediction differ on that input, the bug is real.&lt;/p&gt;

&lt;p&gt;I built a tool that does this and published it under the Apache 2.0 licence. It is called Leafparity. It reads an XGBoost, LightGBM or scikit-learn tree model and its ONNX conversion and returns either a proof of equivalence or the list of disagreements with witness inputs. The examples above are reproduced in its repository, each as a short script you can run yourself.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it does not cover
&lt;/h2&gt;

&lt;p&gt;It reads XGBoost gbtree models with numerical splits, LightGBM gbdt and rf models, and scikit-learn trees and ensembles, compared against ONNX TreeEnsembleRegressor and TreeEnsembleClassifier. It does not handle categorical splits, HistGradientBoosting, XGBoost dart, LightGBM linear trees, or other target formats such as PMML or compiled code. When it meets something it does not support, it says so and refuses to give a verdict instead of guessing.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you run converted tree models
&lt;/h2&gt;

&lt;p&gt;If a converted tree model is in production in your company, I would like to run the check on it. Send the original model and the ONNX file, and you get the report. I am looking for a small number of real models to try it on, and for the first few the check is free.&lt;/p&gt;

&lt;p&gt;Code and examples: &lt;a href="https://github.com/Leafparity/leafparity" rel="noopener noreferrer"&gt;https://github.com/Leafparity/leafparity&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;More about the tool: &lt;a href="https://leafparity.com" rel="noopener noreferrer"&gt;https://leafparity.com&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Milivoje Simonović&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>python</category>
      <category>onnx</category>
      <category>testing</category>
    </item>
  </channel>
</rss>
