DEV Community

Arthur031221
Arthur031221

Posted on

Reading a paper with its source page beside the answer

I often pause on a dense paragraph in a research paper and want a quick explanation. A chat response can help, but I still need to see the sentence and page that supported it. I built Paper Margins as a small local reader for that moment.

The reader opens a local PDF or an arXiv ID. I can select a passage in the rendered page and choose Explain or Translate. The right margin shows the model response and the extracted page text used for that request. Clicking an excerpt returns to the page. I can also ask a question, summarize the current section, or request a short paper summary.

A selected passage, explanation, source excerpts, and citation preview

The demo uses an original two page PDF included in the repository. It cites the Transformer paper so the reference preview can resolve a real arXiv record. The selected text, explanation, and citation lookup in the GIF come from the running tool.

How the reader works

PDF.js renders the pages in the browser and extracts their text on the local Node server. For a question, the server scores short text chunks by words in the question and sends a few matching excerpts to the configured model. For a selected passage, it sends that passage and nearby text from the same page. The response panel shows those excerpts separately from the generated answer, so I can check the answer against what the model actually received.

Ollama is the default model connection. I can also configure a compatible chat endpoint. A local PDF stays on my computer with the default local model. If I set a remote model address, the excerpts used for a request go to that service. Clicking a reference sends its arXiv ID, DOI, or printed reference text to a metadata service; it does not send the complete PDF.

Limits and feedback

This is a reader for one paper at a time. Scanned pages need OCR before text actions work. PDF text order can be wrong in complicated layouts, which affects section and reference detection. Long sections use up to five extracted pages. The excerpts show what the model saw, but they do not prove that its answer is correct.

I would like examples where a two column PDF, a citation format, or a section heading causes a wrong result. The source, tests, sample PDF, and demo are at https://github.com/Arthur031221/paper-margins.

Top comments (0)