DEV Community

Solon Framework
Solon Framework

Posted on

One Interface, Seven Formats: How Solon AI Turns Files, Web Pages, and Even Database Schemas into RAG Documents

Every RAG pipeline starts the same way: you have stuff, and the model needs Documents.

The interesting question is how far that idea stretches. Solon AI answers it with a deliberately small contract — and then pushes it across seven formats, including one you probably haven't tried feeding to a retriever: your database schema.

This is a source-code tour of solon-ai-rag-loaders. All claims below are checked against the current source tree; where a class behaves in a way you wouldn't guess from its name, I'll point it out.

The Contract Is Three Methods

public interface DocumentLoader {
    DocumentLoader additionalMetadata(String key, Object value);
    DocumentLoader additionalMetadata(Map<String, Object> metadata);
    List<Document> load() throws IOException;
}
Enter fullscreen mode Exit fullscreen mode

That's the entire contract: two metadata methods and one load(). No provider field, no API key, no vendor. Seven Maven sub-modules implement it (solon-ai-load-markdown, -pdf, -word, -excel, -html, -ppt, -ddl), each pulling only its own parsing dependency — commonmark, PDFBox, POI, jsoup, Tika.

The base class AbstractOptionsDocumentLoader adds the options pattern with two entry points:

MarkdownLoader loader = new MarkdownLoader(file)
        .options(o -> o.codeBlockAsNew(true));
// or, if you already hold an Options instance:
loader.options(myOptions);
Enter fullscreen mode Exit fullscreen mode

A SupplierEx<InputStream> constructor appears in every loader, so your source can be a file, a URL, a byte array, or anything else that can produce a stream lazily.

The Default Splitting Tells You What Each Format Means

Seven loaders, and no single "chunk size" knob. Instead, each loader picks its default unit of meaning — and the defaults disagree on purpose:

Loader Default unit Default mode
MarkdownLoader Section (per heading) AST walk, headings always split
PdfLoader Page LoadMode.PAGE
WordLoader Paragraph LoadMode.PARAGRAPH
PptLoader Whole document LoadMode.SINGLE
ExcelLoader Sheet, batched at 200 rows JSON rows
HtmlSimpleLoader Whole page Single document
DdlLoader Table One SHOW CREATE TABLE each

That asymmetry is the design. A paragraph is the natural retrieval unit for prose; a page is the natural unit for a PDF; a slide deck usually makes more sense as one document; a table is a complete thought. You can override the defaults (PdfLoader goes SINGLE, WordLoader goes SINGLE, PptLoader splits on "\n\n\n"), but the out-of-the-box behavior already encodes a per-format answer to "what is a chunk here?"

Markdown: Splitting on the AST, Not on Regex

MarkdownLoader doesn't slice text with regexes. It parses the document with commonmark into an AST and walks it with a visitor:

  • Headings always start a new document. Not an option — a rule.
  • Three switches default to off: horizontalLineAsNew, blockquoteAsNew, codeBlockAsNew.
  • Fenced code blocks are more subtle. When codeBlockAsNew(true), the code block starts its own document. Either way, a fenced block always ends its document — so code never bleeds into the prose chunk that follows it.
  • The produced documents carry metadata you can filter on later: category=header_1..6 with a title, category=code_block with lang, or category=blockquote.

One nuance worth knowing before you rely on metadata: the visitor writes title/category onto the current document while walking. If a section has no heading text before its content, the metadata simply won't be there for that chunk. Fine for retrieval; worth remembering if you build UI on top of it.

PDF and Word: The Same Two Ideas, Different Truth

PdfLoader (PDFBox) defaults to one Document per page, each stamped with page, total_pages, and a summary of "Page 3" — handy in a search UI. Switch to LoadMode.SINGLE and you get the whole file as one document, pages joined by "\n\f", with just a pages count.

WordLoader handles both binary eras: it checks the stream with POI's FileMagic and routes .docx (OOXML) and legacy .doc (OLE2) to different readers. It defaults to paragraph mode — one Document per paragraph — with a SINGLE escape hatch.

Excel: Rows in, JSON Out

ExcelLoader (POI + snack4) treats the first non-empty row of a sheet as the header row, then maps every following row to {column: value} and serializes batches as JSON documents. Two defaults shape its behavior:

  • 200 rows per document. A sheet with 620 rows becomes 4 documents. Set documentMaxRows(-1) to keep one document per sheet.
  • An empty row stops the sheet. The read loop breaks, so anything after the first blank row is silently ignored — by design, trailing blank rows shouldn't kill the parse, but data below a blank row won't be indexed. Keep that in mind with hand-edited spreadsheets.

Formula cells are read as their formula text, not computed values.

PowerPoint: Trust Tika

PptLoader doesn't parse slide XML itself. It hands the stream to Apache Tika's AutoDetectParser and gets body text back. Default is SINGLE — the whole deck as one document; PAGE mode splits on "\n\n\n" if your decks have predictable slide breaks.

DDL: Your Schema Is Already a Document

This is the one that changes how you think about the pipeline. DdlLoader connects to a plain DataSource (no ORM, no entities) and emits one Document per table containing its DDL:

DdlLoader loader = new DdlLoader(dataSource);   // MySQL config built in
loader.options(o -> o.loadOptions("shop", null)); // schema only: all its tables
List<Document> docs = loader.load();
Enter fullscreen mode Exit fullscreen mode

Three granularities via loadOptions(schema, table): whole instance, one schema, one table. The default configuration is MySQL (information_schema + SHOW CREATE TABLE, system schemas excluded), but every SQL string is a template — the loader runs them through Solon's own expression engine (SnEL.evalTmpl), so you can rewire it for another database by replacing the template set in a DdlLoadConfig.

One detail I like: SHOW CREATE TABLE returns CREATE TABLE \order(...) — a table name that's only meaningful inside its schema. The loader rewrites the header to CREATE TABLE \shop.\order(...) so every retrieved DDL document is self-describing, and stamps metadata("table", "order") so your filter layer can target tables directly.

The use case writes itself: point it at production (read-only!), and your AI assistant retrieves schema facts instead of hallucinating column names.

After load(): One Shape Downstream

Whatever the format, load() hands you List<Document> — content plus metadata plus the fluent fields (title, url, summary, id, embedding, score). From here, everything is format-agnostic: embed, store in a Repository, attach as a tool. The loaders are the only place in the pipeline where format-specific knowledge lives.

When You Might Skip This Module

To be fair to your architecture review: if your corpus is already clean Markdown, you might not need seven loaders — Solon AI's splitter story covers embedding-time splitting separately. The loaders earn their keep when sources are heterogeneous (office files, web pages, live schema) or when the natural unit (page, paragraph, table) should decide the chunk, not a character count.

Wrapping Up

solon-ai-rag-loaders is a good example of a small contract held firmly: three methods, seven implementations, and per-format defaults that encode real opinions instead of one generic knob. The DDL loader alone is worth a look if you build assistants that need to talk about your database accurately.

Top comments (2)

Collapse
 
agentsearchhq profile image
AgentSearch •

Nice tour of the loaders, the jsoup-based HTML one is where I've been bitten most. One cheap guard that helped: check the extracted text length before chunking, and if it comes back short or empty (typical for JS-rendered pages), route that URL to a headless render step instead of indexing nav and footer boilerplate. It keeps a surprising amount of junk out of the vector store.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.