I write decoders for WebSocket payloads. Someone's app sends 214 bytes, I have no schema, and I have to guess what they are.
The obvious way to build this is a chain of attempts: try JSON, try base64, try gzip, try MessagePack, return whatever sticks. That's what I built first. It was worse than useless, and fixing it changed how I think about this kind of code.
Wrong answers cost more than no answer
MessagePack is where I learned this. It's self-describing — the first byte tells you the type — so a decoder can start reading almost any byte sequence and produce something. Feed it a chunk of an encrypted blob and it will hand you back a map with two keys and an array, confidently.
You then spend twenty minutes trying to work out what that structure means in the app you're debugging. It means nothing. It's noise that happened to start with a valid type byte.
The fix is one rule: a decoder either consumes the entire payload or it fails. No leftover bytes, no partial match, no "close enough". If 200 of 214 bytes parse and 14 are left over, that's a rejection, not a result.
function decode(bytes) {
const { value, consumed } = tryMessagePack(bytes)
if (consumed !== bytes.length) throw new Error('partial')
return value
}
Boring rule, big effect. False positives dropped to roughly zero, and I stopped debugging my own decoder's imagination.
There's a second-order version of the same problem. A Thrift Compact payload without a schema decodes fine structurally — you get field 1, field 2, field 3 with their types. But the field names live in the schema you don't have, so what you're looking at is f1: 8172, f2: "…". That's honest, and occasionally useful, and it's important the UI doesn't dress it up as more than it is.
Some things can't be decoded, and that's not a bug
Discord's gateway taught me this one.
It can compress with zlib-stream. Not "each message is zlib-compressed" — a single zlib stream, opened once when the connection opens, with every message appended to it. Each message ends with 00 00 FF FF, the Z_SYNC_FLUSH marker.
The consequence took me a while to accept: you cannot decompress message 400 on its own. zlib builds a dictionary from everything that came before. Message 400 is compressed against messages 1 through 399. Without them you don't have partial information, you have none.
So the decoder needs a persistent decompression context per connection, fed every frame in order, from the start. Mine keeps one DecompressionStream('deflate') alive per connection, with a queue so frames can't overtake each other, and correlates output back to input frames by the flush boundaries.
And if you attach in the middle of a conversation, it returns nothing. Not corrupted output — nothing. I spent a while looking for the bug before accepting there isn't one. The bytes genuinely don't contain the message.
That limitation is worth surfacing to the user rather than hiding, because it tells them something actionable: reload the page with the tool already running, and it works.
It's also the strongest argument I know for capturing from page load instead of from the moment someone opens DevTools. Stream compression makes late attachment useless, and the handshake is where the interesting failures live anyway.
Detecting what you can't read
If a decoder is going to give up honestly, it needs to know the difference between "I don't have a decoder for this" and "nobody has a decoder for this, it's encrypted".
The obvious measurement is entropy. Encrypted data should look random, compressed and encoded data less so. I implemented it and it doesn't work.
Measured on real payloads:
- encrypted blob: 5.3 bits/char
- base64-encoded JSON: 5.1 bits/char
Those distributions overlap. Pick a threshold anywhere in that range and you'll mislabel readable data as encrypted, or the reverse, depending on which way you lean. Base64 is just dense enough to look like ciphertext by that measure.
What works is dumber. Base64-decode the string, then ask what fraction of the resulting bytes are printable characters:
function looksOpaque(bytes) {
let printable = 0
for (const b of bytes) {
if (b === 9 || b === 10 || b === 13 || (b >= 32 && b < 127)) printable++
}
return printable / bytes.length < 0.75
}
Under about 75% printable and it's encrypted or otherwise opaque. Decoded JSON, decoded text, decoded XML — all overwhelmingly printable. Ciphertext isn't. Two clean populations, no overlap in practice.
A caveat: this runs after the real decoders, not before. gzip and LZ4 payloads aren't printable either, and you don't want to flag as encrypted something you can actually decompress. Hex strings and hashes need excluding too — they're printable, but they decode to nothing meaningful.
The shape that shows up most often in the encrypted bucket looks like this:
[0x01][key id][ciphertext…]
A version byte, a key identifier, then bytes. Once you've seen it a few times you recognise it instantly, which is itself useful — it means stop, this is not a puzzle.
Giving up is a feature
The UI paints those fields red, with a tooltip that says "likely encrypted / opaque binary". It decodes nothing. It's one of the more useful things in the tool.
Not because it tells you what the data is. Because it tells you to stop looking. Knowing a field is permanently unreadable turns a twenty-minute rabbit hole into a two-second decision to go check something else.
Three rules, then, for this kind of code:
Consume everything or fail. Don't hide a limitation that has a workaround the user could act on. And when you can't read something, say so, instead of producing a structure that looks like an answer.
I build this into Wirepeek, a free DevTools panel for WebSocket traffic — disclosure, it's mine. But the rules are general, and I arrived at all three by getting them wrong first.
Top comments (0)