<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: César Freitas</title>
    <description>The latest articles on DEV Community by César Freitas (@shumai).</description>
    <link>https://dev.to/shumai</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4149032%2Fd8a94861-4f35-4eb4-b60a-3dc87a3b6f65.png</url>
      <title>DEV Community: César Freitas</title>
      <link>https://dev.to/shumai</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/shumai"/>
    <language>en</language>
    <item>
      <title>Turn any PDF into clean Markdown for your RAG pipeline</title>
      <dc:creator>César Freitas</dc:creator>
      <pubDate>Tue, 29 Sep 2026 09:57:14 +0000</pubDate>
      <link>https://dev.to/shumai/turn-any-pdf-into-clean-markdown-for-your-rag-pipeline-3ak2</link>
      <guid>https://dev.to/shumai/turn-any-pdf-into-clean-markdown-for-your-rag-pipeline-3ak2</guid>
      <description>&lt;p&gt;Every RAG project hits the same wall on day two: the PDFs.&lt;/p&gt;

&lt;p&gt;You have a folder of them — contracts, manuals, invoices, scanned reports — and you need the text inside. So you reach for a PDF library, run it, and get back a wall of characters with the tables flattened into gibberish, the headings gone, and the two-column pages interleaved line by line. Feed that to an embedding model and your retrieval quality is ruined before you even start.&lt;/p&gt;

&lt;p&gt;Then you hit the scanned ones. No text layer at all. Now you're wiring up an OCR engine, tuning it, and gluing the two paths together.&lt;/p&gt;

&lt;p&gt;I got tired of rebuilding this for every project, so I turned it into an API. Send a PDF, get back clean Markdown that keeps its structure.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "clean" actually means
&lt;/h2&gt;

&lt;p&gt;The goal isn't just extracting characters — it's preserving the shape of the document so a language model can use it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Headings stay headings, so you can chunk on them.&lt;/li&gt;
&lt;li&gt;Tables come back as Markdown tables, not scrambled rows.&lt;/li&gt;
&lt;li&gt;Lists stay lists.&lt;/li&gt;
&lt;li&gt;Reading order is respected on multi-column pages.&lt;/li&gt;
&lt;li&gt;Scanned pages are OCR'd automatically, in the same call.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last point matters: you send the file and you don't care whether it's digital or a photo of a page. The API figures it out.&lt;/p&gt;

&lt;h2&gt;
  
  
  The two endpoints
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;pdf-to-markdown&lt;/strong&gt; — the workhorse. PDF in (as a file or a URL), clean Markdown out, plus the page count and whether OCR kicked in. This is what you chunk and embed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;extract&lt;/strong&gt; — for the structured case. Point it at an invoice or a receipt and get back fields — totals, dates, vendor, line items — as JSON, so you skip the regex entirely.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why an API instead of a local library
&lt;/h2&gt;

&lt;p&gt;Two reasons. First, the "digital PDF plus scanned PDF plus tables plus reading order" combination is genuinely fiddly to get right, and it's the kind of plumbing that has nothing to do with your actual product. Second, it's deterministic and stateless: same file in, same Markdown out, in one request, with nothing to install or keep updated.&lt;/p&gt;

&lt;p&gt;If your RAG retrieval is only as good as your ingestion — and it is — this is the cheapest place to buy a big quality jump.&lt;/p&gt;

&lt;p&gt;It's live on RapidAPI with a free tier: &lt;a href="https://rapidapi.com/cesaricf79/api/document-intelligence3" rel="noopener noreferrer"&gt;https://rapidapi.com/cesaricf79/api/document-intelligence3&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;How are you handling PDF ingestion in your RAG stack right now? I'd love to hear what's working and what still hurts.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>llm</category>
      <category>rag</category>
    </item>
    <item>
      <title>Stop wrestling with broken JSON from your LLM</title>
      <dc:creator>César Freitas</dc:creator>
      <pubDate>Tue, 29 Sep 2026 09:13:51 +0000</pubDate>
      <link>https://dev.to/shumai/stop-wrestling-with-broken-json-from-your-llm-3a4h</link>
      <guid>https://dev.to/shumai/stop-wrestling-with-broken-json-from-your-llm-3a4h</guid>
      <description>&lt;p&gt;If you've shipped anything on top of an LLM, you know this pain: you ask the model for JSON, and you get back JSON wrapped in a code fence, with a friendly "Sure! Here's your data:" in front, single quotes instead of double, a trailing comma, and — on a bad day — the last brace missing because the response got cut off.&lt;/p&gt;

&lt;p&gt;So you write a regex. Then another. Then a try/except that strips fences. Then you discover the model sometimes emits True instead of true. Six months later your "JSON cleaner" is 200 lines of defensive string surgery that everyone is afraid to touch.&lt;/p&gt;

&lt;p&gt;I got tired of copy-pasting that file between projects, so I turned it into a small API. Here's what it handles and how it works.&lt;/p&gt;

&lt;h2&gt;
  
  
  The failure modes
&lt;/h2&gt;

&lt;p&gt;Real LLM output that should be JSON tends to break in a predictable handful of ways:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Markdown code fences around the JSON.&lt;/li&gt;
&lt;li&gt;Prose around it — "Here is the result: { ... } Hope this helps!"&lt;/li&gt;
&lt;li&gt;Single quotes instead of double.&lt;/li&gt;
&lt;li&gt;Unquoted keys.&lt;/li&gt;
&lt;li&gt;Trailing commas.&lt;/li&gt;
&lt;li&gt;Python/JS literals — None, True, False, NaN, undefined.&lt;/li&gt;
&lt;li&gt;Inline comments.&lt;/li&gt;
&lt;li&gt;Truncation — the response hit the token limit and just... stops.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The first seven are cosmetic. The eighth is the nasty one: you have to walk the string, track the open braces and brackets and whether you're inside a string, and close everything back up in the right order.&lt;/p&gt;

&lt;h2&gt;
  
  
  One call
&lt;/h2&gt;

&lt;p&gt;Send the broken string; get back valid parsed JSON, a normalized text version, and the list of fixes that were applied. The truncated case is closed automatically — a fragment that starts an object with an unclosed array comes back as a complete, valid object.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why deterministic, not "ask another model to fix it"
&lt;/h2&gt;

&lt;p&gt;You could send the broken JSON to a second LLM call and ask it to fix it. But that's slower, costs tokens, is non-deterministic, and can hallucinate values that were never there. Repairing JSON is a parsing problem, not a reasoning problem — so it should be solved with a parser. Same input, same output, every time, in milliseconds.&lt;/p&gt;

&lt;h2&gt;
  
  
  It's part of a small toolbox
&lt;/h2&gt;

&lt;p&gt;The same API also has the endpoints I kept needing next to JSON repair:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;context-compress — trim text to a token budget before sending it to a model (removes HTML and boilerplate, dedupes, cuts on sentence boundaries).&lt;/li&gt;
&lt;li&gt;srt-clean — turn SRT/VTT captions into clean text for prompting, and tell you how many tokens you saved.&lt;/li&gt;
&lt;li&gt;log-triage — paste raw Docker/systemd/stack-trace logs and get the exception type, message, and likely cause.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If this saves you from maintaining yet another clean_json.py, it's live on RapidAPI with a free tier: &lt;a href="https://rapidapi.com/cesaricf79/api/llm-dev-utilities" rel="noopener noreferrer"&gt;https://rapidapi.com/cesaricf79/api/llm-dev-utilities&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;What's your worst "the model returned almost-JSON" story? I'm curious how many of these failure modes I'm still missing.&lt;/p&gt;

</description>
      <category>api</category>
      <category>llm</category>
      <category>programming</category>
      <category>softwaredevelopment</category>
    </item>
  </channel>
</rss>
