DEV Community

Cover image for Dear JavaScript developers
Vishal Pandey
Vishal Pandey

Posted on

Dear JavaScript developers

One Package to Parse Them All — Meet pageslice

As a JavaScript developer, I was tired of installing a different package for every file type I needed to parse. One for PDF. One for DOCX. One for Excel. One for CSV. Each with its own API, its own quirks, and its own set of headaches.

Why not just use one package that handles all of it — cleanly?

That's exactly why I built pageslice.

What Can pageslice Do?

Right now, pageslice can parse:

  • PDF files
  • DOCX files
  • XLSX / XLS spreadsheets
  • CSV files
  • TXT / MD files

And we're not stopping here — OCR support and more file formats are on the way.

But Wait, There's More

You'd think a document parser is all pageslice does?

Nope.

pageslice also comes with built-in LLM support. No need to set up an AI client separately — just bring your API key and pageslice can:

  • Extract structured data using Zod schemas
  • Answer questions directly from your documents
  • Work with any OpenAI-compatible provider (OpenAI, Gemini, OpenRouter, and more)

One package. Parsing + AI. Done.

Let's Talk Speed

Here's how fast pageslice runs on real documents:

Benchmark results

Want to test it on your own machine? Just run:

npx pageslice test
Enter fullscreen mode Exit fullscreen mode

Try It Out

If pageslice saved you some headaches, consider giving it a ⭐ on GitHub — it means a lot!

pageslice 📄⚡

Zero-glue document parser & AI extraction toolkit for Node.js & TypeScript.
Slice any PDF, DOCX, XLSX, TXT, or Markdown into raw text or guaranteed typed JSON using Zod schemas.

npm version License: MIT Zero Native Dependencies


Why pageslice?

Every developer building document AI in JavaScript/TypeScript knows the pain:

  • Installing pdf-parse, mammoth, and fighting node-gyp / C++ native build errors in Next.js, Vercel, and Docker.
  • Writing 80+ lines of glue code to handle file formats, buffers, and clean whitespace.
  • Dealing with LLMs returning broken JSON or missing schema fields.

pageslice eliminates all of that. One single package. Zero native C++ dependencies. Pure JavaScript. Works offline for text parsing, and supports OpenAI, Google Gemini, OpenRouter, or local Ollama for AI.


Key Features

  • 📦 All-in-One: PDF, DOCX, XLSX/XLS (Excel), TXT, MD, CSV, JSON support out of the box.
  • 🔓 0 API Keys Needed for Text: Extract…

Feedback, suggestions, and contributions are always welcome. Thanks for reading! 🙌

Top comments (0)