I was reading a post about a homemade config format (in Russian) where the author explained why the format needs a "take this as is" marker: so that 00544 does not quietly become 544. I went to check my own tool, and it did exactly that.
The tool is datadiff. It compares JSON, YAML, CSV, TOML and XML files by their parsed data instead of their lines, so reordered keys and reformatting produce no noise. For CSV that means deciding which cells are numbers. While fixing the zeros I found three more ways the same few lines lost what was written, all with one cause: Rust's standard library accepts more as a number than a person filling in a spreadsheet means.
CSV has no types
JSON knows that 3 is a number and "3" is a string. CSV does not know anything: every cell is text. But a diff that compares data should not report 100 → 100.0, or 1e3 → 1000, as a change. So datadiff reads a cell that looks like a number as a number. The code did the obvious thing:
fn csv_field_to_value(field: &str) -> Value {
let trimmed = field.trim();
if trimmed.is_empty() {
return Value::String(String::new());
}
if let Ok(i) = trimmed.parse::<i64>() {
return Value::Number(Number::Int(i));
}
if let Ok(f) = trimmed.parse::<f64>() {
return Value::Number(Number::Float(f));
}
Value::String(field.to_string())
}
Try an integer, then a float, otherwise keep the text. Short and wrong.
What parse says yes to
Here is what the two parsers return for a few inputs (rustc 1.98.1):
| input | parse::<f64>() |
parse::<i64>() |
|---|---|---|
"Nan", "nan", "+nan"
|
Ok(NaN) |
Err |
"inf", "Infinity", "INFINITY"
|
Ok(inf) |
Err |
"1e400" |
Ok(inf) |
Err |
"00544" |
Ok(544.0) |
Ok(544) |
"12345678901234567890" |
Ok(1.2345678901234567e19) |
Err(PosOverflow) |
"1_000", "0x10", " 5"
|
Err |
Err |
None of this is a bug in Rust. f64 accepts inf, infinity and nan in any case, an exponent too large for a double rounds to infinity instead of failing, and 00544 is a perfectly valid way to write 544. The problem is what each of these "yes" answers does to data that was never meant as a number.
Four ways to lose what was written
Leading zeros. ZIP codes, article numbers and phone extensions keep their zeros for a reason. 00544 became the integer 544, so changing it to 544 was reported as no change at all. This is the same rule that makes 100 and 100.0 equal, applied to a value where nobody means the number.
Words. A name column with Nan in it became NaN. And NaN is not equal to itself, so an untouched row was reported as modified. The opposite also happened: inf and Infinity are different text, but both became the same infinity and compared equal.
Long integers. i64::MAX is 9223372036854775807, nineteen digits. A twenty-digit account number fails to parse as i64, falls through to f64, and a double has 53 bits of mantissa, about 16 significant decimal digits. 12345678901234567890 and 12345678901234567891 turned into the same float, so the change disappeared.
Overflow. 1e400 became infinity, and so did 1e401, so changing one into the other was reported as no change.
The rule: a number only if nothing is lost
The fix is one rule. A cell becomes a number only when reading it as a number loses nothing that was written. Everything else stays text.
fn csv_field_to_value(field: &str) -> Value {
let trimmed = field.trim();
if trimmed.is_empty() {
return Value::String(String::new());
}
if is_plain_number(trimmed) {
if let Ok(i) = trimmed.parse::<i64>() {
return Value::Number(Number::Int(i));
}
// An integer too long for i64 would lose digits as a float.
let is_integer = trimmed
.trim_start_matches(['+', '-'])
.bytes()
.all(|b| b.is_ascii_digit());
if !is_integer {
if let Ok(f) = trimmed.parse::<f64>() {
if f.is_finite() {
return Value::Number(Number::Float(f));
}
}
}
}
Value::String(field.to_string())
}
/// Whether a CSV cell can be read as a number without losing what was
/// written: only digits, sign, point and exponent (so `NaN` and `inf` stay
/// text), and no leading zero before another digit (so `00544` stays an
/// identifier rather than becoming 544).
fn is_plain_number(s: &str) -> bool {
let numeric_chars = s
.bytes()
.all(|b| b.is_ascii_digit() || matches!(b, b'+' | b'-' | b'.' | b'e' | b'E'));
let unsigned = s.trim_start_matches(['+', '-']).as_bytes();
let leading_zero = unsigned.len() > 1 && unsigned[0] == b'0' && unsigned[1].is_ascii_digit();
numeric_chars && !leading_zero
}
Each check closes one of the four holes. The character whitelist keeps Nan and inf as text before parse ever sees them. The leading-zero check keeps 00544 an identifier, while 0.5 and 0 still pass. An integer that does not fit in i64 stays text instead of falling through to a float. And is_finite catches 1e400.
Long integers stay text rather than going into an i128. A twenty-digit value in a CSV is usually an identifier, such as an account, a card or an order number, and an identifier is compared exactly as written.
The tests are black-box: write two small CSV files, run the binary, check the exit code (0 means no differences, 1 means differences).
#[test]
fn csv_words_are_not_parsed_as_special_floats() {
// "Nan" is a name, not NaN: identical cells must not differ.
let out = csv_cell_change("Nan", "Nan");
assert_eq!(out.status.code(), Some(0), "{:?}", out);
// "inf" and "Infinity" are different text, not the same infinity.
let out = csv_cell_change("inf", "Infinity");
assert_eq!(out.status.code(), Some(1), "{:?}", out);
}
The others cover 00544/544, -007/-7 and 0123.5/123.5 (must differ), the twenty-digit account numbers (must differ), and 100/100.0, 0.5/0.50, 1e3/1000 (must stay equal).
What is still not right
The rule fixed what I had seen, not everything. Running the current binary on a few more pairs:
| old | new | result |
|---|---|---|
9007199254740993 |
9007199254740992 |
changed (both fit in i64, compared exactly) |
0.1 |
0.10000000000000001 |
no change: both are the same double |
+5 |
5 |
no change |
+12345678901234567890 |
12345678901234567890 |
changed: long integers are compared as text |
Decimals are still doubles, so two values that differ past the sixteenth digit compare equal. And the boundary of i64 is visible: a leading plus sign is ignored for short integers and counts as a difference for long ones. Both could be fixed without floats and without a new dependency, by normalizing the decimal string (sign, digits, exponent) and comparing that. If you have solved this more cleanly, I would like to hear how.
The fix shipped in datadiff 0.4.1. datadiff is a CLI that diffs JSON, YAML, CSV, TOML and XML by their parsed data and plugs into git diff; the source and the tests above are at github.com/dimanovikov/datadiff.
I wrote this post with the help of an AI assistant. The code, numbers and outputs come from the datadiff repository and rustc 1.98.
Top comments (0)