I build BCD, a hosted remote MCP server that lets ChatGPT and Claude read files and run commands on your own computer. Our logs showed file reads failing far more often on Windows than anywhere else. There were three causes, and you'd never see any of them testing on a Mac.
1. pdf.js and a trailing \
We passed pdf.js (pdfjs-dist) its standard font folder like this:
const standardFontDataUrl =
path.resolve(here, '../node_modules/pdfjs-dist/standard_fonts') + path.sep;
On macOS and Linux that ends in /. On Windows path.sep is \, so it ends in \. pdf.js treats the value as a URL and checks the last character:
function getFactoryUrlProp(val) {
if (typeof val !== "string") return null;
if (val.endsWith("/")) return val;
throw new Error(`Invalid factory url: "${val}" must include trailing slash.`);
}
So on Windows, getDocument() threw for every PDF, whatever its content, and from the outside it just looked like "this PDF can't be read". The fix is one line:
const fontFolderUrl = (folder, separator = path.sep) =>
folder.split(separator).join('/') + '/';
In Node, pdf.js appends the font file name and hands the result to fs.readFile. Node's fs on Windows reads C:/Users/.../standard_fonts/FoxitSans.pfb with forward slashes just fine. We added a unit test that feeds in a Windows-style path, and a step in our Windows CI smoke test that reads a real PDF.
Lesson: when you append path.sep, check whether the value is used as a path or as a URL.
2. PowerShell writes UTF-16
In Windows PowerShell 5.1, command > out.txt writes UTF-16LE with a byte order mark (FF FE). Our text reader was UTF-8 only and treated any NUL byte as binary. UTF-16 has a NUL after every ASCII character, so those files were all rejected. An AI writing output to a file with PowerShell and then reading it back is a common flow, so this hurt.
Now the first two bytes pick the decoder: new TextDecoder('utf-16le', { fatal: true }) (or utf-16be). When we read the end of a large file, cutting in the middle can split a character, so the window starts after the next line break (0A 00 in UTF-16LE).
3. Excel CSVs in CP949, and Node's euc-kr trap
On Korean Windows, Excel saves CSV in CP949. There's a trap here: Node's TextDecoder('euc-kr') is not CP949. It can't decode the extended Hangul syllables, and it doesn't throw. It silently produces control characters:
new TextDecoder('euc-kr', { fatal: true }).decode(Buffer.from([0x8c, 0x63]))
// '\x8Cc' (wrong, and no exception)
iconv.decode(Buffer.from([0x8c, 0x63]), 'cp949')
// '똠'
What we do now:
- If UTF-8 decoding hits an invalid sequence, read the file again in the legacy code page for the computer's language (CP949 for Korean, CP932 for Japanese, CP936 or CP950 for Chinese) with iconv-lite.
- If that result contains
U+FFFD, it isn't text. - Files with NUL bytes are real binaries and are never retried.
- The read reports which encoding it used (
encoding: "cp949"), so the AI knows before it writes the file back.
Bonus: classify failures honestly
A file that really isn't text is now recorded as an invalid argument, not a tool failure. If your own bugs and your users' inputs land in the same bucket, you won't see the next bug.
Disclosure: I'm the maker of BCD. It lets ChatGPT and Claude work on your own Mac, Windows or Linux computer: files, commands, documents and a browser.
Top comments (0)