DEV Community

zhihu wu
zhihu wu

Posted on

Your Base64 Decoding Failed Because of Whitespace (and 3 Other Silent Traps)

You paste a Base64 string into a decoder, and it throws InvalidCharacterError or returns garbage. Nine times out of ten the alphabet is fine — it's invisible whitespace, padding, or the URL-safe variant that broke it.

Here are the four traps I hit most often, and how to spot each one in seconds.

1. Line breaks (the MIME 76-column wrap)

Email-era encoders still wrap output at 76 characters with \r\n. Python's base64.encodebytes() does this, and so do PEM/JWKS files. atob() in the browser refuses any whitespace, so a string that looks perfect silently fails. Strip all whitespace before decoding, and prefer base64.b64encode() over encodebytes() when you control the encoder.

2. Missing or excess padding

= pads the final group to a multiple of 4 characters. YWJjZA (6 chars) and YWJjZA== are the same bytes, but many strict decoders reject the unpadded form. Some SDKs strip padding; others require it. If a decoder fails, add = until length % 4 == 0.

3. base64url vs base64

JWTs, URL query params, and filenames use the URL-safe alphabet: - and _ replace + and /, and padding is usually dropped. Feed a JWT segment straight into a standard decoder and it breaks. In Python: base64.urlsafe_b64decode(s + "=" * (-len(s) % 4)). In JS, do the character swap yourself — replace(/-/g, '+').replace(/_/g, '/') — before calling atob.

4. Encoding text is not encoding bytes

atob() returns a binary string, so UTF-8 text comes back as mojibake. Decode to bytes first, then to text: new TextDecoder().decode(Uint8Array.from(atob(s), c => c.charCodeAt(0))), or base64.b64decode(s).decode('utf-8') in Python.

The five-second check

Before you debug your code, look at the string itself: any whitespace, any - or _, does the length divide by 4? That diagnoses most "corrupted Base64" bugs without touching the source.

When I need to eyeball a payload — checking padding, spotting a base64url segment inside a JWT, or confirming a decode round-trips my UTF-8 — I use CodeToolbox's Base64 encoder/decoder. It runs entirely in the browser (nothing is uploaded), accepts both alphabets, and shows character counts so you can see the padding at a glance.

Top comments (0)