The scikit-learn ONNX converter documentation has a tutorial about what happens when you switch a decision tree from float64 to float32. On its test set, the original model and its ONNX copy differ by at most about 156. I took that same example and asked a different question: what is the largest difference that can exist anywhere in the input space? The answer is 556. If missing values are allowed, it is 721.
The test set was not wrong. It was just never going to find the places where the two models disagree.
Why a test set cannot see it
A tree model is a set of boxes. Every input falls into exactly one leaf per tree, and the leaf decides the output. When you convert a model to ONNX, the thresholds that define those boxes can change slightly. A threshold stored as a float64 in the original becomes a float32 in the converted file, so it moves by a tiny amount. Inputs that fall between the old and the new threshold now go down the other branch.
Those inputs live in a sliver a few billionths wide next to each threshold. A random sample of a few thousand test rows will almost never land there. Production traffic will eventually, because real data has prices, ratios and counts that sit exactly on round numbers, and round numbers are where thresholds like to be.
When an input does land in a sliver, the output does not change by a rounding error. It changes by a whole leaf value, which is where the 156 and the 556 come from.
Three places the differences hide
Threshold rounding. The float32 cast described above. It affects models whose thresholds are float64 in the original and float32 after conversion, such as LightGBM and scikit-learn trees.
Zero treated as missing. LightGBM can be trained with zero_as_missing=True, which means an input of exactly 0.0 is handled as a missing value. When I converted such a model with onnxmltools, the converter ignored that setting. An input of exactly zero then took a different path in every tree, and the raw scores differed by up to about 295.
NaN routing. Since scikit-learn 1.3, trees can route NaN to the left or the right child at each node, learned during training. In the converted model, the NaN went the other way at the affected nodes.
None of these are exotic. Each one comes from a reasonable converter doing a reasonable thing in a place where the two libraries have different rules.
Checking it without sampling
Because both models are finite sets of boxes, there is a way to compare them without guessing. Walk the two models together over the whole input space. At every region, ask whether the original and the converted model reach corresponding leaves. If they do everywhere, the models agree for every possible input, up to floating point rounding of the outputs. If they do not, you can list each region where they disagree, the exact set of inputs affected, and the largest difference.
For every problem found, you also want a concrete input you can run through both real runtimes, so nobody has to trust the analysis. If the XGBoost, LightGBM or scikit-learn prediction and the onnxruntime prediction differ on that input, the bug is real.
I built a tool that does this and published it under the Apache 2.0 licence. It is called Leafparity. It reads an XGBoost, LightGBM or scikit-learn tree model and its ONNX conversion and returns either a proof of equivalence or the list of disagreements with witness inputs. The examples above are reproduced in its repository, each as a short script you can run yourself.
What it does not cover
It reads XGBoost gbtree models with numerical splits, LightGBM gbdt and rf models, and scikit-learn trees and ensembles, compared against ONNX TreeEnsembleRegressor and TreeEnsembleClassifier. It does not handle categorical splits, HistGradientBoosting, XGBoost dart, LightGBM linear trees, or other target formats such as PMML or compiled code. When it meets something it does not support, it says so and refuses to give a verdict instead of guessing.
If you run converted tree models
If a converted tree model is in production in your company, I would like to run the check on it. Send the original model and the ONNX file, and you get the report. I am looking for a small number of real models to try it on, and for the first few the check is free.
Code and examples: https://github.com/Leafparity/leafparity
More about the tool: https://leafparity.com
Milivoje Simonović
Top comments (0)