A large language model only knows what it saw during training. Ask it about your internal wiki, last week's release notes, or a customer's support history, and it will either shrug or — worse — confidently make something up. Retrieval-Augmented Generation (RAG) is the standard fix: before the model answers, you retrieve the relevant facts from your own data and hand them over as context.
Most RAG tutorials are written in Python. This one is in Java, using Solon AI — the AI module of the Solon framework. We'll go from a pile of raw text to a grounded answer in a single, runnable file, then look at how to swap the in-memory store for Redis, load real documents, and filter by metadata.
Every API in this article was checked against the solon-ai source, so the method names are the real ones — not the hallucinated ones.
The moving parts
A RAG pipeline in Solon AI is built from five small pieces:
| Piece | Type | Job |
|---|---|---|
EmbeddingModel |
model | turn text into vectors |
Document |
data | a chunk of content + metadata + score |
DocumentSplitter |
util | slice long text into chunks |
Repository |
store | save vectors, search by similarity |
ChatModel |
model | generate the final answer |
The flow is always the same: split → embed → store, then at query time embed the question → search → augment the prompt → generate.
Add the dependency
The solon-ai aggregate pulls in the core plus the OpenAI / Ollama / DashScope / Gemini / Anthropic dialects, which is everything you need for a minimal pipeline.
<dependency>
<groupId>org.noear</groupId>
<artifactId>solon-ai</artifactId>
<!-- replace with the latest stable release from Maven Central -->
<version>4.1.0</version>
</dependency>
EmbeddingModel, ChatModel, InMemoryRepository, Document, and the built-in splitters all live in solon-ai-core, so a minimal RAG needs no extra vector-store or loader modules.
The whole pipeline in one file
Here is an end-to-end example. It uses a local Ollama server so you can run it without any API keys — point the URLs and models at OpenAI or DashScope if you prefer.
import org.noear.solon.ai.chat.ChatModel;
import org.noear.solon.ai.chat.ChatResponse;
import org.noear.solon.ai.chat.message.ChatMessage;
import org.noear.solon.ai.embedding.EmbeddingModel;
import org.noear.solon.ai.rag.Document;
import org.noear.solon.ai.rag.RepositoryStorable;
import org.noear.solon.ai.rag.repository.InMemoryRepository;
import org.noear.solon.ai.rag.splitter.SplitterPipeline;
import org.noear.solon.ai.rag.splitter.TokenSizeTextSplitter;
import org.noear.solon.ai.rag.util.QueryCondition;
import java.util.Arrays;
import java.util.List;
public class MiniRag {
public static void main(String[] args) throws Exception {
// 1) Embedding model: turns text into vectors
EmbeddingModel embeddingModel = EmbeddingModel.of("http://127.0.0.1:11434/api/embed")
.provider("ollama") // or "openai" / "dashscope"
.model("bge-m3")
.batchSize(10)
.build();
// 2) In-memory vector store (must be given an EmbeddingModel)
RepositoryStorable repository = new InMemoryRepository(embeddingModel);
// 3) Your source content
String rawText = "Solon is a Java application development framework. "
+ "It offers its own IoC/AOP container, a lightweight web layer, "
+ "and Solon AI for building LLM applications. "
+ "Solon starts fast and has a small memory footprint, "
+ "which makes it a good fit for GraalVM native images.";
List<Document> rawDocs = Arrays.asList(
new Document(rawText).title("About Solon").url("https://solon.noear.org")
);
// 4) Split long text into chunks (default chunkSize = 500 tokens)
SplitterPipeline pipeline = new SplitterPipeline()
.next(new TokenSizeTextSplitter(500));
List<Document> chunks = pipeline.split(rawDocs);
// 5) Store — save() embeds each chunk in batches automatically
repository.save(chunks);
// 6) Retrieve
String question = "What is Solon and why is it good for native images?";
QueryCondition condition = new QueryCondition(question)
.limit(4) // default is 4
.similarityThreshold(0.4D); // default is 0.4
List<Document> hits = repository.search(condition);
// 7) Chat model for generation
ChatModel chatModel = ChatModel.of("http://127.0.0.1:11434/api/chat")
.provider("ollama") // or "openai" / "dashscope"
.model("qwen2.5")
.build();
// 8) Augment the prompt with retrieved context, then ask
ChatMessage userMsg = ChatMessage.ofUserAugment(question, hits);
ChatResponse resp = chatModel.prompt(userMsg).call();
if (resp.getError() != null) {
throw resp.getError();
}
System.out.println(resp.getMessage().getContent());
}
}
That's the entire pipeline. Let's unpack the parts that matter.
What the API is actually doing
Document — plain data, no magic factory
A Document holds an id, content, a Map<String, Object> metadata, a transient score (filled in during search), and a float[] embedding. There is no Document.of(...) factory — you use the constructor and chain setters:
new Document("some content")
.title("About Solon")
.url("https://solon.noear.org")
.metadata("category", "framework");
Splitting is a pipeline
SplitterPipeline lets you chain splitters with next(...). The built-in TokenSizeTextSplitter cuts on token count (default 500, backed by jtokkit's CL100K_BASE), and RegexTextSplitter cuts on a pattern (default \n\n). Chain them when you want "split on blank lines, then cap each piece at N tokens".
save() embeds for you
RepositoryStorable.save(...) is the write path — note the name is save, not insert or store. Internally it batches by embeddingModel.batchSize() and calls embed(...) for each batch, so you never touch vectors by hand. There's also asyncSave(...), deleteById(String...), and existsById(String).
QueryCondition controls retrieval
The constructor takes the query string; everything else is chained:
-
limit(int)— how many chunks to return (default 4) -
similarityThreshold(double)— minimum score to keep (default 0.4) -
filterExpression(String)— metadata filter (more on this below)
Augmentation: two ways
The interesting one is ChatMessage.ofUserAugment(question, context). It wraps your question and the retrieved documents into a single user message using this template:
{question}
Now: {current time}
References: {context}
So the model gets the question, the current time (handy for "latest" style questions), and the references — all in one shot.
If you'd rather not orchestrate the search yourself, Repository has a one-liner that does search + wrap together:
// search(question) + ofUserAugment(...) in a single call
ChatMessage augmented = repository.promptAugment(question);
ChatResponse resp = chatModel.prompt(augmented).call();
System.out.println(resp.getContent());
Going beyond in-memory
Swap in a real vector store
InMemoryRepository is great for demos, but for production you'll want a persistent store. Solon AI ships repositories for Redis, Milvus, Qdrant, pgvector, Elasticsearch, OpenSearch, Chroma, Weaviate, MySQL, MariaDB, and more — each is a separate Maven module (solon-ai-repo-redis, solon-ai-repo-milvus, …).
Redis, for example, uses a builder that takes the embedding model plus a Jedis client:
// dependency: org.noear:solon-ai-repo-redis
RedisRepository repository = RedisRepository
.builder(embeddingModel, jedisClient)
.indexName("my_docs")
.build();
The Repository interface is identical across stores, so everything downstream — save, search, QueryCondition — stays exactly the same. Swapping the store never touches your retrieval code.
Load real documents
For real content you rarely start from a String. The TextLoader (from File, URI, URL, byte[], or a stream) lives in the core; separate modules add MarkdownLoader, PdfLoader, HtmlSimpleLoader (note: not HtmlLoader), WordLoader, ExcelLoader, and PptLoader. Every loader returns List<Document>, so it drops straight into the splitter → store flow.
Filter by metadata with SnEL
QueryCondition.filterExpression(String) accepts a SnEL expression and narrows the search to documents whose metadata matches — vector similarity and a structured filter, in one query:
QueryCondition condition = new QueryCondition(question)
.limit(4)
.filterExpression("category == 'framework' AND year >= 2024");
Under the hood the string is parsed by SnEL.parse into a portable expression tree, and each vector store rewrites that tree into its own native filter syntax. (If you want the full story on how one SnEL expression targets Redis, Milvus, Qdrant, and pgvector, I wrote about that in an earlier post.)
Let the agent retrieve on its own
Everything above is "retrieve first, then ask". Solon AI also supports agentic RAG via RepositoryTool, which wraps a Repository as a callable tool (@ToolMapping("repository_query")). Hand it to a ChatModel and the model decides when to search — useful for multi-turn conversations where not every question needs a lookup.
ChatModel chatModel = ChatModel.of(apiUrl)
.provider("openai")
.model("gpt-4o-mini")
.defaultToolAdd(new RepositoryTool(repository))
.build();
Wrapping up
The mental model is small: split → embed → store, then embed → search → augment → generate. Solon AI gives you each step as a plain, composable piece — Document, DocumentSplitter, Repository, EmbeddingModel, ChatModel — with no framework ceremony. Start with InMemoryRepository to prove the flow, switch to Redis or pgvector when you need persistence, and reach for filterExpression and RepositoryTool when your retrieval gets more demanding.
Same interface all the way down, no rewrite when you scale up. That's the part I like.
Project: solon.noear.org · Source: github.com/opensolon/solon-ai
Top comments (0)