<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Dima Novikov</title>
    <description>The latest articles on DEV Community by Dima Novikov (@dimanovikov).</description>
    <link>https://dev.to/dimanovikov</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4120988%2F6ea4c0ea-f963-4d25-9b1f-1c9f9986c73a.jpg</url>
      <title>DEV Community: Dima Novikov</title>
      <link>https://dev.to/dimanovikov</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/dimanovikov"/>
    <language>en</language>
    <item>
      <title>Why datadiff matches arrays by key instead of computing tree edit distance</title>
      <dc:creator>Dima Novikov</dc:creator>
      <pubDate>Sat, 19 Sep 2026 12:39:14 +0000</pubDate>
      <link>https://dev.to/dimanovikov/why-datadiff-matches-arrays-by-key-instead-of-computing-tree-edit-distance-2ehp</link>
      <guid>https://dev.to/dimanovikov/why-datadiff-matches-arrays-by-key-instead-of-computing-tree-edit-distance-2ehp</guid>
      <description>&lt;p&gt;I wrote &lt;a href="https://github.com/dimanovikov/datadiff" rel="noopener noreferrer"&gt;datadiff&lt;/a&gt; to compare structured data (JSON, YAML, CSV, TOML, XML) based on parsed values rather than line differences. Two files are parsed into a single internal tree representation, and the tool prints the differences as property paths—for example, &lt;code&gt;spec.replicas: 3 → 5&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The main architectural decision was made upfront: I intentionally did not implement an optimal tree edit distance algorithm.&lt;/p&gt;

&lt;p&gt;The idea came from a Hacker News discussion about Graphtage, Trail of Bits' semantic diff tool. Graphtage finds the minimal edit script between two trees, which is the mathematically correct way to describe what changed. However, most comments in that thread focused on execution time: diffing two Kubernetes pods took several minutes, and a 45 KB CSV was projected to take days. Computing tree edit distance is computationally heavy, and for configuration files, the precision rarely justifies the runtime.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cost of tree edit distance
&lt;/h2&gt;

&lt;p&gt;Handling objects in a structured diff is straightforward because keys provide identity. Arrays are different: elements have no inherent keys, so the algorithm has to determine which old element corresponds to which new one. An optimal diff evaluates potential pairings to find the sequence of edits with the lowest total cost. This allows it to detect when an element has simply moved rather than being deleted and recreated elsewhere, but it also causes runtime to scale poorly.&lt;/p&gt;

&lt;p&gt;To see the practical impact, I ran a benchmark on an Apple M4 comparing Graphtage 0.5.0 with datadiff 0.4.1. The input files contained JSON arrays of objects with seven fields (including one nested object), where the updated file had its rows shuffled and 1% of the email fields modified:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;objects&lt;/th&gt;
&lt;th&gt;file size&lt;/th&gt;
&lt;th&gt;datadiff &lt;code&gt;--key id&lt;/code&gt;
&lt;/th&gt;
&lt;th&gt;Graphtage&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;14 KB&lt;/td&gt;
&lt;td&gt;0.005 s&lt;/td&gt;
&lt;td&gt;1.76 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;250&lt;/td&gt;
&lt;td&gt;36 KB&lt;/td&gt;
&lt;td&gt;0.005 s&lt;/td&gt;
&lt;td&gt;9.38 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;500&lt;/td&gt;
&lt;td&gt;73 KB&lt;/td&gt;
&lt;td&gt;0.005 s&lt;/td&gt;
&lt;td&gt;38.1 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1,000&lt;/td&gt;
&lt;td&gt;146 KB&lt;/td&gt;
&lt;td&gt;0.010 s&lt;/td&gt;
&lt;td&gt;timeout (&amp;gt;120 s)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;10,000&lt;/td&gt;
&lt;td&gt;1.5 MB&lt;/td&gt;
&lt;td&gt;0.076 s&lt;/td&gt;
&lt;td&gt;skipped&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;100,000&lt;/td&gt;
&lt;td&gt;15 MB&lt;/td&gt;
&lt;td&gt;0.46 s&lt;/td&gt;
&lt;td&gt;skipped&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Graphtage scales at roughly O(n²) or worse on this input. To be clear, the comparison is asymmetric: Graphtage solves the general problem without knowing how to identify items, whereas datadiff relies on being told that &lt;code&gt;id&lt;/code&gt; is the primary key. In practice, however, config files almost always contain natural identifiers—such as container names, user IDs, or table primary keys.&lt;/p&gt;

&lt;h2&gt;
  
  
  How datadiff handles matching
&lt;/h2&gt;

&lt;p&gt;The internal value representation is standard:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="k"&gt;pub&lt;/span&gt; &lt;span class="k"&gt;enum&lt;/span&gt; &lt;span class="n"&gt;Value&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;Null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nf"&gt;Bool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;bool&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="nf"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="nf"&gt;String&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;String&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="nf"&gt;Array&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;Vec&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;Value&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="nf"&gt;Object&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BTreeMap&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nb"&gt;String&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Value&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Using &lt;code&gt;BTreeMap&lt;/code&gt; for objects normalizes key ordering during parsing. Two objects with identical fields in different orders produce identical maps, so reformatting never triggers a diff. It also guarantees deterministic iteration order, which is convenient when using the tool as a git textconv driver.&lt;/p&gt;

&lt;p&gt;Objects are compared by iterating over both maps. Arrays are compared positionally by default, unless the user provides a &lt;code&gt;--key&lt;/code&gt; argument. When a key is specified, datadiff constructs an index map for each array:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="nf"&gt;key_map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;arr&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;Value&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;Option&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;BTreeMap&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nb"&gt;String&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;usize&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="k"&gt;mut&lt;/span&gt; &lt;span class="n"&gt;map&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nn"&gt;BTreeMap&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;new&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="n"&gt;arr&lt;/span&gt;&lt;span class="nf"&gt;.iter&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="nf"&gt;.enumerate&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nn"&gt;Value&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;Object&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;obj&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nb"&gt;None&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="p"&gt;};&lt;/span&gt;
        &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;kv&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;obj&lt;/span&gt;&lt;span class="nf"&gt;.get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;?&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;scalar_key_text&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;kv&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;?&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;map&lt;/span&gt;&lt;span class="nf"&gt;.insert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="nf"&gt;.is_some&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nb"&gt;None&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;// duplicate key value&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="nf"&gt;Some&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;map&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If any element is not an object, is missing the specified key, or has a non-scalar key value, or if two elements share a key value, the function returns &lt;code&gt;None&lt;/code&gt; and the tool falls back to positional index matching. This fallback behavior is deliberate: &lt;code&gt;--key name&lt;/code&gt; applies globally across the entire document, even though many nested arrays are simple string lists that lack a &lt;code&gt;name&lt;/code&gt; field.&lt;/p&gt;

&lt;p&gt;When both maps build successfully, matching elements is a direct lookup. Matching objects are diffed recursively, missing items are reported as insertions or deletions, and the key is incorporated into the output path: &lt;code&gt;containers[name=api].image&lt;/code&gt; instead of &lt;code&gt;containers[1].image&lt;/code&gt;. The total time complexity is O(n log n) due to map construction.&lt;/p&gt;

&lt;p&gt;The limitation of this approach is obvious: without &lt;code&gt;--key&lt;/code&gt;, reordered arrays produce positional diffs. datadiff will not detect moved items unless explicitly told how to identify them, and it never reports renames. For source code analysis, this would be insufficient; for configuration management, it avoids unbounded execution times.&lt;/p&gt;

&lt;h2&gt;
  
  
  Number equality and CSV edge cases
&lt;/h2&gt;

&lt;p&gt;In a configuration diff, &lt;code&gt;1&lt;/code&gt; and &lt;code&gt;1.0&lt;/code&gt; should compare equal, but serde parses them into distinct internal types. datadiff handles equality across representations explicitly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="k"&gt;pub&lt;/span&gt; &lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="nf"&gt;numbers_eq&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;match&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nn"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;Int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nn"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;Int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nn"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;UInt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nn"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;UInt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nn"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;Int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nn"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;UInt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="nb"&gt;u64&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nn"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;UInt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nn"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;Int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;y&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;y&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="nb"&gt;u64&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="k"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="nf"&gt;.as_f64&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="nf"&gt;.as_f64&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;CSV files introduced additional edge cases because they have no schema or types. My initial implementation attempted to parse fields as &lt;code&gt;i64&lt;/code&gt;, then &lt;code&gt;f64&lt;/code&gt;, and fell back to raw strings. This failed in three distinct ways:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;"00544".parse::&amp;lt;i64&amp;gt;()&lt;/code&gt; parses to 544, masking changes where leading zeros were stripped (such as postal codes).&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;f64::from_str&lt;/code&gt; accepts &lt;code&gt;"nan"&lt;/code&gt;, &lt;code&gt;"inf"&lt;/code&gt;, and &lt;code&gt;"infinity"&lt;/code&gt; case-insensitively. A string value &lt;code&gt;"Nan"&lt;/code&gt; parsed as &lt;code&gt;NaN&lt;/code&gt;, and since &lt;code&gt;NaN != NaN&lt;/code&gt;, an unchanged row produced a false diff.&lt;/li&gt;
&lt;li&gt;20-digit identifiers overflow &lt;code&gt;i64&lt;/code&gt;, parse as &lt;code&gt;f64&lt;/code&gt;, lose lower-order bits, and compare equal to different numeric values.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;To resolve this, strings are only converted to numbers if the conversion does not alter the underlying text:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="nf"&gt;is_plain_number&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;numeric_chars&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;
        &lt;span class="nf"&gt;.bytes&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="nf"&gt;.all&lt;/span&gt;&lt;span class="p"&gt;(|&lt;/span&gt;&lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;|&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="nf"&gt;.is_ascii_digit&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;||&lt;/span&gt; &lt;span class="nd"&gt;matches!&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sc"&gt;b'+'&lt;/span&gt; &lt;span class="p"&gt;|&lt;/span&gt; &lt;span class="sc"&gt;b'-'&lt;/span&gt; &lt;span class="p"&gt;|&lt;/span&gt; &lt;span class="sc"&gt;b'.'&lt;/span&gt; &lt;span class="p"&gt;|&lt;/span&gt; &lt;span class="sc"&gt;b'e'&lt;/span&gt; &lt;span class="p"&gt;|&lt;/span&gt; &lt;span class="sc"&gt;b'E'&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;unsigned&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="nf"&gt;.trim_start_matches&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="sc"&gt;'+'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sc"&gt;'-'&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;&lt;span class="nf"&gt;.as_bytes&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;leading_zero&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;unsigned&lt;/span&gt;&lt;span class="nf"&gt;.len&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="n"&gt;unsigned&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sc"&gt;b'0'&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="n"&gt;unsigned&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="nf"&gt;.is_ascii_digit&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="n"&gt;numeric_chars&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="n"&gt;leading_zero&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Values with leading zeros, words such as &lt;code&gt;nan&lt;/code&gt; or &lt;code&gt;inf&lt;/code&gt;, and integers too long for &lt;code&gt;i64&lt;/code&gt; now remain strings, while &lt;code&gt;100&lt;/code&gt; and &lt;code&gt;100.0&lt;/code&gt; still compare equal.&lt;/p&gt;

&lt;h2&gt;
  
  
  A macOS test race condition
&lt;/h2&gt;

&lt;p&gt;During integration testing, test cases generated temporary directories named using &lt;code&gt;std::process::id()&lt;/code&gt; and &lt;code&gt;SystemTime::now()&lt;/code&gt; nanoseconds. After I added a test that calls this helper in a loop, tests occasionally failed because one test read fixture files written by another test running in parallel in the same process.&lt;/p&gt;

&lt;p&gt;On macOS, &lt;code&gt;SystemTime&lt;/code&gt; updates at microsecond resolution rather than nanoseconds. Concurrent tests started at roughly the same time were generating identical paths and overwriting each other's test fixtures. Appending an &lt;code&gt;AtomicUsize&lt;/code&gt; counter to the directory names resolved the collision.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the tradeoff bought
&lt;/h2&gt;

&lt;p&gt;Matching arrays by key is a weaker guarantee than an optimal edit script.&lt;br&gt;
datadiff cannot tell you that an element moved unless the key says so, and it&lt;br&gt;
has to be told which field the key is. What it gets back is the case&lt;br&gt;
configuration files actually hit: a 15 MB file diffed in half a second instead&lt;br&gt;
of not finishing at all.&lt;/p&gt;

&lt;p&gt;The code is at &lt;a href="https://github.com/dimanovikov/datadiff" rel="noopener noreferrer"&gt;github.com/dimanovikov/datadiff&lt;/a&gt;&lt;br&gt;
under MIT / Apache-2.0, and the benchmark above is&lt;br&gt;
&lt;a href="https://github.com/dimanovikov/datadiff/blob/main/bench/compare_graphtage.py" rel="noopener noreferrer"&gt;&lt;code&gt;bench/compare_graphtage.py&lt;/code&gt;&lt;/a&gt;&lt;br&gt;
in the same repository.&lt;/p&gt;

</description>
      <category>rust</category>
      <category>algorithms</category>
      <category>performance</category>
      <category>cli</category>
    </item>
    <item>
      <title>Teaching git diff to read JSON and YAML</title>
      <dc:creator>Dima Novikov</dc:creator>
      <pubDate>Fri, 11 Sep 2026 13:25:01 +0000</pubDate>
      <link>https://dev.to/dimanovikov/teaching-git-diff-to-read-json-and-yaml-12h1</link>
      <guid>https://dev.to/dimanovikov/teaching-git-diff-to-read-json-and-yaml-12h1</guid>
      <description>&lt;p&gt;A formatter reordered the keys in a Kubernetes manifest and &lt;code&gt;git diff&lt;/code&gt; lit up&lt;br&gt;
six lines. One of them changed the replica count. I only noticed after the&lt;br&gt;
deploy.&lt;/p&gt;

&lt;p&gt;That is not a git bug. A line diff compares lines, and in structured data the&lt;br&gt;
line is the wrong unit: move a key two lines up and nothing changed, but every&lt;br&gt;
line did. So you learn to read config diffs slowly, which means sometimes you&lt;br&gt;
read them fast.&lt;/p&gt;

&lt;p&gt;I wrote a tool for this, and the part worth writing about is not the diffing.&lt;br&gt;
It is getting it inside &lt;code&gt;git diff&lt;/code&gt;, because git has two extension points for&lt;br&gt;
this and they behave differently in ways the docs do not spell out.&lt;/p&gt;
&lt;h2&gt;
  
  
  The two hooks
&lt;/h2&gt;

&lt;p&gt;Git lets you replace a diff in two places.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;An external diff driver&lt;/strong&gt; takes over entirely. Git hands your program both&lt;br&gt;
versions and prints nothing of its own. You own the output.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git config diff.datadiff.command &lt;span class="s2"&gt;"datadiff git-diff"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;A textconv filter&lt;/strong&gt; is narrower. Your program converts each version to text,&lt;br&gt;
and git line-diffs the two outputs as usual.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git config diff.datadiff.textconv &lt;span class="s2"&gt;"datadiff normalize"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then point file types at the driver:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s1"&gt;'*.yaml diff=datadiff'&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; .gitattributes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The surprise is that these cover different commands. I expected the external&lt;br&gt;
driver to apply everywhere. It does not.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Command&lt;/th&gt;
&lt;th&gt;What runs&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;git diff&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;external driver&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;git log -p&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;textconv&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;git show&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;textconv&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;git blame&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;textconv&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Git only calls an external driver for &lt;code&gt;git diff&lt;/code&gt;. History commands fall back to&lt;br&gt;
a line diff, and textconv is your only way in there. Which turns out to be the&lt;br&gt;
right split anyway: when I ask "what did I just change" I want data paths, and&lt;br&gt;
when I read history I want a normal diff, just without the noise.&lt;/p&gt;

&lt;p&gt;So &lt;code&gt;git diff&lt;/code&gt; now says:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;~ spec.replicas: 3 → 5
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and &lt;code&gt;git log -p&lt;/code&gt; says:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt; apiVersion: apps/v1
 spec:
&lt;span class="gd"&gt;-  replicas: 3
&lt;/span&gt;&lt;span class="gi"&gt;+  replicas: 5
&lt;/span&gt;   template:
     image: app:1.0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The reordering is gone from the second one because both sides get canonicalised&lt;br&gt;
before git compares them. Same file, same commit, two different questions.&lt;/p&gt;
&lt;h2&gt;
  
  
  Four things that bit me
&lt;/h2&gt;
&lt;h3&gt;
  
  
  Git passes seven arguments, not two
&lt;/h3&gt;

&lt;p&gt;An external driver is called as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;path old-file old-hex old-mode new-file new-hex new-mode
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The two versions are arguments two and five. Argument one is the path. A tool&lt;br&gt;
that takes &lt;code&gt;old new&lt;/code&gt; cannot be plugged in directly, it will compare the&lt;br&gt;
filename against the old version. You either wrap it in a shell function that&lt;br&gt;
picks out &lt;code&gt;$2&lt;/code&gt; and &lt;code&gt;$5&lt;/code&gt;, or you teach the tool git's convention. I did the&lt;br&gt;
second, the shim has problems I will get to.&lt;/p&gt;
&lt;h3&gt;
  
  
  A non-zero exit kills the whole diff
&lt;/h3&gt;

&lt;p&gt;This one cost me the most. My tool exits 1 when it finds differences, which is&lt;br&gt;
correct for a CLI and fatal for a diff driver. Git reads any non-zero status as&lt;br&gt;
the driver having failed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;fatal: external diff died, stopping at deploy.yaml
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Not "this file failed". The entire diff stops, and every file after that one is&lt;br&gt;
never shown.&lt;/p&gt;

&lt;p&gt;It gets worse. The obvious case is a file that is not structured data at all,&lt;br&gt;
which you can avoid with a careful &lt;code&gt;.gitattributes&lt;/code&gt;. The case you cannot avoid&lt;br&gt;
is a file that is invalid &lt;strong&gt;right now&lt;/strong&gt; because you are in the middle of editing&lt;br&gt;
it. You type half a JSON object, run &lt;code&gt;git diff&lt;/code&gt; to see what you have done, and&lt;br&gt;
git blows up. A driver that does that gets uninstalled the same day.&lt;/p&gt;

&lt;p&gt;So the driver has to swallow its own errors:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;deploy.json
  no semantic diff: invalid JSON: EOF while parsing a value at line 2
  run `git diff --no-ext-diff` to see this file as plain text
cfg.toml
~ x.n: 1 → 2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Prints a note, keeps walking, exits zero.&lt;/p&gt;

&lt;h3&gt;
  
  
  The format has to come from argument one
&lt;/h3&gt;

&lt;p&gt;Git stages the two versions in temporary files. In my testing the basename&lt;br&gt;
survived (&lt;code&gt;/tmp/git-blob-CVrzDV/conf.json&lt;/code&gt;), so sniffing the extension off the&lt;br&gt;
temp file happens to work. I would not rely on it. Argument one is the real&lt;br&gt;
path and it is always there, so detect the format from that and fall back to&lt;br&gt;
the temp files, not the other way around.&lt;/p&gt;

&lt;p&gt;This is the first reason to prefer a subcommand over a shell shim: &lt;code&gt;$2&lt;/code&gt; and&lt;br&gt;
&lt;code&gt;$5&lt;/code&gt; are all the shim has, so it is guessing.&lt;/p&gt;
&lt;h3&gt;
  
  
  Read bytes, not text, in the fallback
&lt;/h3&gt;

&lt;p&gt;My textconv mode passes a file through unchanged when it cannot parse it. The&lt;br&gt;
first version did this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="nn"&gt;std&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nn"&gt;fs&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;read_to_string&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;file&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="nf"&gt;.unwrap_or_default&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Which is fine until someone has a config in CP1251. &lt;code&gt;read_to_string&lt;/code&gt; fails on&lt;br&gt;
invalid UTF-8, &lt;code&gt;unwrap_or_default&lt;/code&gt; hands back an empty string, both sides&lt;br&gt;
normalize to nothing, and git reports the file as unchanged. The file vanished&lt;br&gt;
from &lt;code&gt;git log -p&lt;/code&gt; entirely. That is worse than no integration: a plain diff at&lt;br&gt;
least tells you the versions differ.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;raw&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nn"&gt;std&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nn"&gt;fs&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;file&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="nf"&gt;.unwrap_or_default&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="nn"&gt;std&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nn"&gt;io&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;stdout&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="nf"&gt;.write_all&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;raw&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="nf"&gt;.ok&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Copy bytes. A fallback that loses data is not a fallback.&lt;/p&gt;

&lt;h2&gt;
  
  
  The shim, and why I stopped using it
&lt;/h2&gt;

&lt;p&gt;You can get most of this with one line and no new code:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git config diff.datadiff.command &lt;span class="s1"&gt;'f() { datadiff --exit-zero "$2" "$5"; }; f'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It works. I shipped that first and documented it. Four things pushed me to a&lt;br&gt;
real subcommand:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;it cannot name the file it is diffing, and git prints nothing around driver
output, so a diff across five files is unreadable&lt;/li&gt;
&lt;li&gt;the format is guessed from temp files&lt;/li&gt;
&lt;li&gt;a file that is invalid mid-edit aborts everything&lt;/li&gt;
&lt;li&gt;the function syntax needs a POSIX shell, so Windows users are out&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you are wiring up an existing tool, start with the shim. If you own the&lt;br&gt;
tool, spend the afternoon.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this does not fix
&lt;/h2&gt;

&lt;p&gt;Submodules and symlinks reach the driver if your &lt;code&gt;.gitattributes&lt;/code&gt; pattern is&lt;br&gt;
broad enough, and there is nothing useful to say about either, so they land in&lt;br&gt;
the same note-and-continue path. With the narrow patterns you actually want&lt;br&gt;
(&lt;code&gt;*.json&lt;/code&gt;, &lt;code&gt;*.yaml&lt;/code&gt;) they never get there.&lt;/p&gt;

&lt;p&gt;Colour is handled for you, incidentally. Git pipes driver output to a pager, so&lt;br&gt;
a library that checks for a tty turns colour off on its own.&lt;/p&gt;




&lt;p&gt;The tool is &lt;a href="https://github.com/dimanovikov/datadiff" rel="noopener noreferrer"&gt;datadiff&lt;/a&gt;, Rust,&lt;br&gt;
MIT/Apache-2.0, and it also does CSV, TOML and XML plus a &lt;code&gt;--fail-on&lt;/code&gt; mode for&lt;br&gt;
CI. But the git mechanics above are not specific to it. If you maintain&lt;br&gt;
anything that understands a file format better than &lt;code&gt;diff&lt;/code&gt; does, these are the&lt;br&gt;
two hooks and those are the four traps.&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>devops</category>
      <category>git</category>
      <category>rust</category>
    </item>
  </channel>
</rss>
