--
title: "Parsing tweets.js Without the Encoding Headaches"
description: "The X archive file is JavaScript with a JSON array inside it. Strip the wrapper, then deal with the layer that actually breaks things: code points, surrogate pairs and mixed-script counting."
tags: ["javascript", "privacy", "twitter", "unicode"]
canonical_url: https://digital-footprint-health.shop/blog/tweets-js-emoji-unicode-parsing
Unzip an X archive and you get a folder with a handful of data files in it. One of them is tweets.js. The first attempt to read it fails in the same way almost every time, with a parser error on the opening characters.
The file is not damaged. It is just JavaScript rather than JSON, and the array you want is sitting inside it.
The array is wrapped in an assignment
tweets.js starts with a line that assigns a bracketed array to a namespaced variable and ends with a semicolon. The array on its own is valid JSON. Everything around it is not, which is why a strict parser refuses the whole file.
You have three ways past it, and the right one depends on how many files you are handling.
Open it in a text editor, find the first opening bracket, delete everything ahead of it, save as .json. Fine for one file, tedious at ten.
Read the whole string in code, slice from the first opening bracket to the last closing bracket, and hand the slice to the parser.
Write one function and loop over the folder. An archive ships posts, likes and messages with the same wrapper, so this is the version that survives a real export.
The batch route costs about ten extra minutes and removes the step where mistakes happen.
Three symptoms, three different layers
When characters come out wrong, the symptom narrows the search faster than swapping libraries and hoping.
| Symptom | What it usually is | What to do |
|---|---|---|
| CJK text renders as mojibake | Wrong codec on read, commonly UTF-8 decoded as Latin-1 | Specify UTF-8 on read and write; re-extract if bad text was already saved back |
| Emoji render as empty boxes | The terminal or font cannot draw them; the data is intact | Check code points before changing any code |
| Length does not match what you see | Code points, UTF-16 units and bytes are being treated as one number | Fix the counting rule once, at the top of the script |
The second row catches people out most often. Write the string into an HTML file and open it in a browser. If the emoji appear there, the file was never the problem.
Code points, UTF-16 units and bytes
One emoji is one symbol on screen. In UTF-16 it can occupy two units. As UTF-8 bytes it can take four. Measure the same string three ways and you get three different numbers.
The damage lands on truncation. Slice a string by length and you can cut an emoji in half, leaving two halves that render as boxes. To truncate by visible character, expand to code points first and join afterwards.
A check that takes thirty seconds
Print the code point of every character in the string you suspect and read the output. Ordinary text lands in familiar ranges. An isolated value between 55296 and 57343 means a surrogate pair was split somewhere upstream, and the fix belongs in the slicing code rather than in the parser.
CJK text needs a different counting rule
Chinese and Japanese do not separate words with spaces, so splitting on whitespace does nothing useful. Keyword statistics on CJK content want character-level counting or a tokenizer.
Mixed-script posts add a second wrinkle. Half-width and full-width punctuation both show up, and naive matching quietly skips part of the set. Normalising punctuation before comparison costs little and catches a surprising number of misses.
Errors and what they mean
| Error or symptom | Fix |
|---|---|
| Unexpected token in JSON | The prefix was not fully stripped; slice from the first opening bracket |
| Unexpected end of JSON input | A trailing semicolon or extra characters got captured; end at the last closing bracket |
| Parses cleanly but fields are empty | The archive is sharded; the payload sits one array level deeper |
| Emoji count comes out low | Counting by length; switch to code points |
Write the counting rule down
Two people can process the same archive and report different totals, and the gap is almost always definitional. Whether reposts count, whether replies count, whether characters are measured by code point. Put those three answers in a comment at the top of the script and every later comparison becomes meaningful.
Choosing what to write out
Structured data is a starting point. The next need is usually classification, by year, by keyword, or by how much risk a post carries, and all three want a local index so you can query the archive without touching the network. That is the main argument for keeping archive analysis on your own machine, and the trade-offs are set out in local versus cloud processing.
Three output shapes cover most cases. A tabular export filters easily by year or keyword in a spreadsheet and suits manual review. The original nested structure is best for further processing. A local index suits repeated queries across fields.
If the archive is being kept, store at least one copy in the original nested shape. Field definitions have changed between archive versions, and raw data lets you recalculate later rather than re-export.
When fields contain commas
A tabular export breaks the moment a post contains a comma or a line break, because both split columns. Quote every field and escape the contents instead of joining on plain commas. The failure is silent and usually shows up after the import, once the columns are already misaligned.
Handling archives from different years
Archive formats have been revised more than once, so exports from different periods are not identical in shape. Older versions shard posts across several files and some field names have been renamed since.
When one script has to handle archives from several years, run a field probe first. Read the file, print the field names on the top-level objects, compare against the expected list. Failing loudly beats writing empty values with no warning.
Once the index exists, deletion has a basis. Which posts stay and which go should follow rules you set rather than a single date cut-off, and the batch cleanup walkthrough sets out a workflow that keeps a reversible record at each step.
digital-footprint-health.shop runs a free footprint check on your public posts. To see which ones deserve attention, start from the homepage. Supported archive formats are listed on the upload page, and plans sit on the pricing page.
Top comments (0)