DEV Community

Solon Framework
Solon Framework

Posted on

One Interface, Three Search Backends: Web Search Is Just Another Repository

If you build anything with LLMs, you will need web search sooner or later. The model's training data is frozen; your users' questions are not.

In Solon AI, web search is not a special subsystem bolted onto the side. It is a Repository — the same interface a vector store implements, the same interface an in-memory document list implements. That single design decision is the whole story of the solon-ai-rag-searchs family, and it composes in ways you might not expect.

The one method that matters

public interface Repository {
    default List<Document> search(String query) throws IOException {
        return search(new QueryCondition(query));
    }

    List<Document> search(QueryCondition condition) throws IOException;

    default ChatMessage promptAugment(String query) throws IOException {
        List<Document> context = search(query);
        return ChatMessage.ofUserAugment(query, context);
    }
}
Enter fullscreen mode Exit fullscreen mode

One required method. Two defaults on top of it. The interface has been marked @Preview since it landed in 3.1 — which reads as an honest signal: the surface is small enough to grow carefully.

Everything a caller can express goes through QueryCondition:

Field Default Notes
limit 4 max documents back
similarityThreshold 0.4 used by the optional refilter pass
freshness null ONE_DAY / ONE_WEEK / ONE_MONTH / ONE_YEAR
filterExpression null an Expression<Boolean>, parsed from SnEL if you pass a string
searchType VECTOR FULLTEXT, HYBRID (3.3+) for repositories that support them
disableRefilter false skip the similarity re-filter pass

Note what is not here: no provider, no apiKey, no vendor field. Those live in the implementation's builder. The interface never learns which search engine answered the question.

Three implementations, three philosophies

Bocha: the reference implementation

BochaWebSearchRepository repo = BochaWebSearchRepository
        .of("https://api.bochaai.com/v1/web-search")
        .apiKey("sk-...")
        .build();

List<Document> docs = repo.search(new QueryCondition("solon framework").limit(5));
Enter fullscreen mode Exit fullscreen mode

The Bocha adapter (@since 3.1) is deliberately plain: build a JSON body with query, count, and optional freshness, POST it, map data.webPages.value[] into Documents. If the response code is not 200, it throws an IOException — no empty-list disguise, no silent fallback.

There is a comment in the source that tells you why it stays this plain: "此示例,可作为对接其它搜索的参考" — "this example can serve as a reference for integrating other search providers". It is the one you copy when you wire up your own provider.

Baidu AI Search: two modes, one flag

BaiduWebSearchRepository repo = BaiduWebSearchRepository
        .ofAI()                              // or ofBasic()
        .apiKey("bce-...")
        .apiUrl("https://qianfan.baidubce.com/v2/ai_search")
        .build();
Enter fullscreen mode Exit fullscreen mode

The Baidu adapter (@since 3.0, built on Baidu's AI Search V2 endpoint) has two personalities:

  • BASIC — classic results list. No model field is sent; the server infers the mode.
  • AI — sends model (default ernie-3.5-8k; the builder javadoc also lists deepseek-r1, deepseek-v3, and the ERNIE 4.0 turbo variants) and gets back an LLM-synthesized answer plus references.

Here is the detail I like most: in AI mode, the synthesized answer is returned as the first Document in the same list — titled Solon AI智能回答, with metadata("type", "ai_answer"). References follow it with metadata("source", "baidu_search"). The caller does not need a second type or a special branch to consume an AI answer; it is just document zero.

Your limit is translated into a resource_type_filter of type web with top_k. An empty or blank query short-circuits to an empty list without a network call; an error code in the response becomes an IOException. If nothing at all parses out, that is an exception too — "no content found" is surfaced, not swallowed.

Tavily: search, extract, crawl, map

Tavily is the maximalist of the three. The solon-ai-search-tavily module (@since 3.9.5) exposes four operations through TavilySimpleSearchRepository:

TavilySimpleSearchRepository repo = TavilySimpleSearchRepository
        .of("tvly-...")
        .build();

// Standard Repository path — returns List<Document>
List<Document> docs = repo.search(new QueryCondition("solon framework").limit(5));

// Full Tavily power — returns TavilySearchResponse
TavilySearchResponse resp = repo.search(new SearchCondition("solon framework")
        .topic("news")                // general / news / finance
        .searchDepth("advanced")      // basic / advanced / fast / ultra-fast
        .timeRange("week")            // day / week / month / year
        .includeAnswer("basic")       // LLM answer, basic / advanced
        .includeFavicon(true)
        .includeDomains(List.of("github.com"))
        .maxResults(10));

String answer = resp.getText();        // the LLM answer, if requested
List<Document> docs2 = resp.toDocuments();
Enter fullscreen mode Exit fullscreen mode

Beyond search there is extract(condition) (pull clean content from specific URLs), crawl(condition) (walk a site with maxDepth 1–5, maxBreadth up to 500, path regex filters), and map(condition) (structure-only: just the URL list, cheap, meant to feed a later extract).

The standard QueryCondition.freshness is faithfully translated — ONE_DAY → "day", ONE_YEAR → "year" — so the vendor-neutral path still gets time filtering.

The module actually contains two repository classes, and their names are a trap for the unwary:

  • TavilyWebSearchRepository — despite the full-sounding name, this is the simplified one: it implements Repository and forwards search to a delegate.
  • TavilySimpleSearchRepository — despite the humble name, this is the complete one: search with full parameters, extract, crawl, map. It also implements Repository.

The simplified one exposes getFullRepository() so you can start narrow and widen later without rebuilding. (If you find older docs mentioning a TavilySearchRepository, that class name does not exist in the source tree — use TavilySimpleSearchRepository.)

And if you want none of the repository abstraction at all, the underlying TavilyClient is public: ClientBuilder.of(apiKey).apiBase(...).timeout(...).build() gives you the raw four operations.

The optional embedding model — and why you usually shouldn't

All three adapters accept an optional EmbeddingModel. When present, the flow after the HTTP call is identical everywhere:

embeddingModel.embed(docs);
float[] queryEmbed = embeddingModel.embed(condition.getQuery());
return SimilarityUtil.refilter(docs.stream()
        .map(doc -> SimilarityUtil.score(doc, queryEmbed)), condition);
Enter fullscreen mode Exit fullscreen mode

Embed every result, embed the query, score, then refilter — which re-ranks and applies your similarityThreshold and limit.

The module README makes the counterintuitive point explicitly: for web search, similarity re-ranking is usually not meaningful. The snippets come back already ranked by the engine's own relevance machinery, and cosine similarity between a short query and a web snippet is a noisy signal at best. So the guidance is: unless you have a concrete reason, do not pass an embedding model to a web-search repository. The parameter exists for the cases where you do.

This is a nice instance of a general principle: a capability that is off by default and documented as "probably skip this" beats a magic pipeline you cannot turn off.

From passive retrieval to an agent that searches

promptAugment is passive RAG: you search, you glue results into the user message, you send. Fine for FAQs, useless when the model needs to decide whether to search, or what to search for.

Since 3.10.1 there is RepositoryTool:

RepositoryTool tool = new RepositoryTool(webSearchRepo);
// optionally: new RepositoryTool(repo, rerankingModel)
Enter fullscreen mode Exit fullscreen mode

It extends the framework's tool provider base and registers itself as a repository_query tool with two parameters: queries (a list of query strings, hard-capped at 5 per call — the source comment says why: to stop a model from fetching ten pages at once and blowing up its own context) and topK (default 3).

For each query it searches, optionally reranks, and formats results as Markdown with title, relevance score, content, and citation URL. One detail worth copying into your own tools: any document's content is truncated at 2000 characters before it enters the tool response, with an explicit ...(内容过长已截断) marker. Context budgets are defended at the tool layer, not trusted to the model's restraint.

Because it takes any Repository, the same tool class works against a vector store, an in-memory repo, or your Bocha/Baidu/Tavily web repo. One tool, whatever knowledge backend you configured.

When not to use this

  • You need domain filtering on the neutral path — includeDomains/excludeDomains exist only on Tavily's SearchCondition; the shared QueryCondition has no such field. For the others, put the filtering in your query or downstream.
  • You expect streaming search results. The interface returns a completed List<Document>; an AI-mode synthesis on Baidu takes however long it takes.
  • You need hybrid search semantics — searchType(HYBRID) in QueryCondition is for repositories that understand it (3.3+). A web adapter ignores what does not apply to it.

The takeaway

solon-ai-rag-searchs is a small family doing exactly one thing well: making web search interchangeable with every other knowledge source behind Repository. The result is that "add live web knowledge to this agent" becomes a one-line wiring change — and the same wiring upgrades from promptAugment to a self-directed RepositoryTool without touching the backend.

Bocha shows the pattern, Baidu folds an LLM answer into the document list, Tavily goes deep on capabilities — and your agent code sees none of the differences.


All API details above were verified against the source of the solon-ai repository (solon-ai-rag-searchs and solon-ai-core); dependencies are managed by the Solon AI BOM, so no per-artifact version is shown. If you try it, the docs live at solon.noear.org.

Top comments (0)