DEV Community

Royal Simpson Pinto
Royal Simpson Pinto

Posted on

Why Your JSON Signatures Break: A Guide to Deterministic Canonical Serialization in Python

I learned this the hard way while building answerproof, a small tool that signs structured answers so anyone can later verify they were not tampered with. The signing worked. The verification worked. And then, intermittently, verification failed on data that was byte-for-byte semantically identical to what I had signed. Nothing was corrupted. Nothing was malicious. The problem was that I signed a Python dict, serialized it one way when signing, and serialized it a slightly different way when verifying.

That is the whole trap of signing JSON. Let me explain why it happens and how to fix it properly.

The core problem: signatures sign bytes, not meaning

A digital signature is computed over a sequence of bytes. HMAC, Ed25519, RSA, all of them take bytes in and produce a signature. They have no idea that {"a":1,"b":2} and {"b": 2, "a": 1} represent the same object. To a signature function those are two completely different inputs, and they produce two completely different results.

So the moment your data structure passes through JSON, you have introduced ambiguity. The same dict can be serialized in many valid ways:

  • Key order can differ. Python dicts preserve insertion order, but a value that survived a round trip through a database, a queue, or another language may come back with keys in a different order.
  • Whitespace can differ. json.dumps defaults to ", " and ": " separators. Another encoder might use no spaces at all.
  • Numbers can differ. 1.0, 1, and 1e0 are all equal as numbers but different as text.
  • Unicode can differ. "café" can be emitted as literal UTF-8 or as the escaped "café".

Every one of these changes the bytes without changing the meaning. If the signer and the verifier disagree on any of them, verification fails.

The fix is canonicalization: agreeing on exactly one byte representation for any given value, so signing and verifying always see the same input.

A canonical serializer in Python

Python's standard library gives you enough control to do this well. Here is the function I settled on:

import json

def canonical_bytes(obj) -> bytes:
    return json.dumps(
        obj,
        sort_keys=True,          # deterministic key order
        separators=(",", ":"),   # no insignificant whitespace
        ensure_ascii=False,      # stable, explicit unicode as UTF-8
        allow_nan=False,         # reject NaN and Infinity outright
    ).encode("utf-8")
Enter fullscreen mode Exit fullscreen mode

Let me justify each argument, because every one of them closes a specific hole.

sort_keys=True sorts object keys by their Unicode code points. This removes the entire class of key-order bugs. It does not matter what order the keys arrived in; they always come out sorted.

separators=(",", ":") strips the insignificant whitespace that json.dumps adds by default. Now there is exactly one comma between items and one colon between key and value, with nothing extra.

ensure_ascii=False is the choice people argue about. When set to True, Python escapes every non-ASCII character to a \uXXXX sequence. When set to False, it emits the raw UTF-8 characters. Both are deterministic on their own. What matters is that the signer and verifier agree, and that you commit to one. I pick UTF-8 output with ensure_ascii=False and then encode to bytes explicitly, so the whole pipeline is UTF-8 end to end. If you invert the choice, you must invert it on both sides.

allow_nan=False makes the encoder raise an error on float('nan') or float('inf'). Those are not valid JSON, and different parsers handle them differently or not at all. It is far better to fail loudly at signing time than to sign something a verifier cannot even parse.

Signing and verifying with the canonical form

With canonicalization in place, the signing code becomes boringly reliable, which is exactly what you want:

import hmac
import hashlib

def sign(obj, key: bytes) -> str:
    payload = canonical_bytes(obj)
    return hmac.new(key, payload, hashlib.sha256).hexdigest()

def verify(obj, key: bytes, signature: str) -> bool:
    expected = sign(obj, key)
    return hmac.compare_digest(expected, signature)
Enter fullscreen mode Exit fullscreen mode

Notice that both sign and verify route through the same canonical_bytes function. That single shared path is the point. There is no way for the two sides to disagree on serialization because there is only one serializer. Also note hmac.compare_digest, which compares in constant time and avoids leaking information through timing.

Here is the behavior that used to break me, now working:

key = b"secret-key-material"

a = {"question": "2+2", "answer": 4, "meta": {"model": "x", "score": 1}}
b = {"answer": 4, "meta": {"score": 1, "model": "x"}, "question": "2+2"}

sig = sign(a, key)
assert verify(b, key, sig)   # different key order, still verifies
Enter fullscreen mode Exit fullscreen mode

The two dicts have different insertion order at both the top level and inside meta. Because sort_keys recurses into nested objects, both canonicalize to the identical byte string, and the signature holds.

The one honest caveat: numbers

Sorted keys and stripped whitespace are fully solved by the standard library. Numbers are not, and I want to be straight about it.

canonical_bytes({"n": 1}) produces {"n":1}, but canonical_bytes({"n": 1.0}) produces {"n":1.0}. Those are different bytes and therefore different signatures, even though 1 and 1.0 are mathematically equal. Python's json module serializes floats using repr, which is stable within CPython but is not guaranteed to match how another language or another JSON library formats the same number. Large integers, floats near the edge of double precision, and trailing-zero decimals are all places where two systems can legitimately produce different text for the same value.

The honest answer is that JSON has no single canonical number format, so you have to constrain your own data. In answerproof I keep it simple by treating the signed structure as authoritative: whatever numeric types go in are the numeric types that get signed and verified, and I do not let values round trip through a system that might retype them. If you need cross-language canonicalization with strict number rules, look at RFC 8785 (JSON Canonicalization Scheme), which specifies an exact number format. That is more machinery than most projects need, but it is the correct reference when you do.

Takeaways

  • Signatures sign bytes, so you must fix the bytes before you sign.
  • sort_keys=True, separators=(",", ":"), ensure_ascii=False, and allow_nan=False give you a deterministic serializer in four arguments.
  • Route both signing and verifying through the exact same serializer function.
  • Numbers are the one soft spot. Control your numeric types, or adopt RFC 8785 if you need strict cross-language guarantees.

If you want to see this wired into a real sign-and-verify flow, the code lives in my project at github.com/AgentPostmortem/answerproof. Borrow the canonical_bytes function, drop it in front of your signer, and one whole category of flaky verification bugs simply goes away.

Top comments (0)