Mojibake is one of those bugs that looks like corrupted data even when the bytes are still fine. A UTF-8 file opened as Windows-1252 can turn é into é; a Big5 file decoded as UTF-8 can become a wall of unrelated symbols. The hard part is that the visible text often no longer tells you which decoder was used. I built this tool around a safer question: which plausible decoding produces text that looks least broken, and how confident should we be?
That is deliberately different from claiming to “detect the encoding” with certainty. The component tries candidates, scores their output, and shows up to five results so a person can verify the language and context.
File Mode Preserves the Original Bytes
When the original file is available, the UI uses FileReader.readAsArrayBuffer() instead of starting with a JavaScript string:
function loadFile(file) {
fileError.value = "";
fileName.value = file.name;
const reader = new FileReader();
reader.onload = (e) => {
fileBuffer.value = e.target.result;
};
reader.onerror = () => {
fileError.value = t("encodingFixer.fileLoadError");
fileBuffer.value = null;
};
reader.readAsArrayBuffer(file);
}
That choice keeps the raw byte sequence available for each decoding attempt. The catalog includes UTF-8, GBK/GB2312, GB18030, Big5, Shift-JIS, EUC-JP, EUC-KR, several ISO-8859 variants, Windows-1250 through Windows-1257, and KOI8-R/U. For each candidate, TextDecoder runs in fatal mode:
const decoder = new TextDecoder(enc, { fatal: true });
const text = decoder.decode(buf);
An encoding that cannot decode the bytes is skipped rather than turned into a misleading string full of replacement characters. Successful candidates are scored and only those above 10 are kept.
A Score Is a Bundle of Imperfect Signals
The scorer samples at most 2,000 characters, counts printable code points, and penalizes ranges that are common signs of a wrong decode:
const sample = text.slice(0, 2000);
let printable = 0;
let suspicious = 0;
for (let i = 0; i < sample.length; i++) {
const cp = sample.charCodeAt(i);
if (cp === 9 || cp === 10 || cp === 13 || (cp >= 32 && cp !== 127)) {
printable++;
}
if (isSuspiciousCodePoint(cp)) suspicious++;
}
let score = Math.round(printableRatio * 100)
- Math.round(suspiciousRatio * 300);
The suspicious ranges include halfwidth Katakana, private-use characters, Yi characters, and isolated Hangul jamo. The score is then clamped between 0 and 99. UTF-8 receives a small bonus when it decodes cleanly, while raw ISO-8859-1 and Windows-1252 output is penalized if it contains many control-range bytes. Replacement characters are also penalized by their ratio; a result with more missing data should not outrank a clean result merely because most of its characters are printable.
This is heuristic ranking, not language understanding. A short identifier, a code file, or an unusual proper noun may look “suspicious” even when it is correct.
Paste Mode Replays a Mojibake Mistake
Once text has been copied, its original bytes are gone. The paste path uses iconv-lite to try a smaller cross-encoding search: encode the visible mojibake as one of four possible source encodings, then decode those bytes as one of four possible target encodings:
for (const sourceEncoding of SOURCE_ENCODINGS) {
const rawBytes = iconv.encode(originalText, sourceEncoding.enc);
for (const targetEncoding of TARGET_ENCODINGS) {
if (sourceEncoding.enc === targetEncoding.enc) continue;
const restoredText = iconv.decode(rawBytes, targetEncoding.enc);
if (!restoredText || restoredText === originalText) continue;
// score and keep the candidate
}
}
The source list is Big5, GBK, Shift-JIS, and Windows-1252. Targets are UTF-8, Big5, GBK, and Shift-JIS. Results with too many synthetic ? characters are discarded because iconv-lite substitutes characters that do not fit a codec with ?; a string of legal question marks can otherwise score deceptively well. When the input contains no literal question marks, those generated marks are displayed as � so partial recovery is visible.
After scoring, the candidates are sorted from highest confidence to lowest and deduplicated. File mode uses the first 100 characters of the preview as its signature; paste mode compares the full restored text. That prevents the interface from filling with the same readable result produced by several source/target combinations, while still preserving different plausible repairs. The first candidate receives a visual highlight, but the copy and download actions remain available for every candidate. A download is always emitted as a UTF-8 text blob, regardless of the encoding label that produced it, which is convenient for saving a repaired result but worth remembering if another system specifically requires Big5 or Shift-JIS bytes.
Honest Boundaries
The original file is the best input. In file mode, previews are capped at 5,000 characters and the UI returns at most five deduplicated candidates. In paste mode, copying already-decoded gibberish can permanently discard byte information, especially for CJK text; trying every combination cannot recreate bytes that no longer exist. A high confidence score is evidence, not proof, so compare a candidate with known words, headers, or a sample from the source system before saving it over the original.
I turned this into a small free tool: Encoding Fixer (Multilingual Mojibake Repair). It is most useful when you can upload the original .txt, .csv, or .log file and keep the recovery process reversible.
Top comments (0)