<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sergey Nikolaev</title>
    <description>The latest articles on DEV Community by Sergey Nikolaev (@sanikolaev).</description>
    <link>https://dev.to/sanikolaev</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F363352%2F6f7a2da7-fa00-47f5-aaca-a007b1d43350.jpeg</url>
      <title>DEV Community: Sergey Nikolaev</title>
      <link>https://dev.to/sanikolaev</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sanikolaev"/>
    <language>en</language>
    <item>
      <title>Closing the KNN rescoring gap between columnar and row-wise storage</title>
      <dc:creator>Sergey Nikolaev</dc:creator>
      <pubDate>Mon, 05 Oct 2026 09:00:00 +0000</pubDate>
      <link>https://dev.to/sanikolaev/closing-the-knn-rescoring-gap-between-columnar-and-row-wise-storage-2bj</link>
      <guid>https://dev.to/sanikolaev/closing-the-knn-rescoring-gap-between-columnar-and-row-wise-storage-2bj</guid>
      <description>&lt;p&gt;Originally published on the &lt;a href="https://mnt.cr/go/ZVCFKA" rel="noopener noreferrer"&gt;Manticore Search website&lt;/a&gt; on October 5, 2026&lt;/p&gt;

&lt;h1&gt;
  
  
  Closing the KNN rescoring gap between columnar and row-wise storage
&lt;/h1&gt;

&lt;p&gt;Why default KNN rescoring makes columnar vector access slower than row-wise storage, and how Manticore closes most of the gap while preserving out-of-core search performance.&lt;/p&gt;

&lt;p&gt;Manticore supports both row-wise and columnar attribute storage. Row-wise storage works well when the &lt;strong&gt;working set fits in memory&lt;/strong&gt;, but columnar storage becomes especially &lt;strong&gt;useful when large datasets exceed RAM&lt;/strong&gt;: queries that need only a few attributes can read and cache mostly the data they actually use.&lt;/p&gt;

&lt;p&gt;There was one important performance gap between the two layouts. During KNN search, Manticore rescored HNSW candidates using their original full-precision vectors. With the previous columnar access path, that step was much slower than with row-wise storage. In our DBpedia benchmark, row-wise storage delivered 2.80x–3.53x the KNN throughput.&lt;/p&gt;

&lt;p&gt;The problem wasn't the columnar layout itself. Columnar vectors were read through a reusable buffer, which prevented Manticore from keeping stable pointers to multiple vectors and processing them in batches.&lt;/p&gt;

&lt;p&gt;We changed columnar access to use memory mapping. This gives vectors stable addresses while still allowing the operating system to load and evict file pages on demand. The result is 2.57x–3.13x higher KNN throughput, reaching 85–92% of row-wise performance, without sacrificing columnar storage's ability to work efficiently when the data is larger than available memory.&lt;/p&gt;

&lt;h2&gt;
  
  
  How slow can rescoring from columnar storage be?
&lt;/h2&gt;

&lt;p&gt;We measured KNN throughput on a 16-core AMD Ryzen 9 5950X using the DBpedia dataset:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;975,000 vectors&lt;/li&gt;
&lt;li&gt;1,536 coordinates per vector&lt;/li&gt;
&lt;li&gt;1-bit quantization&lt;/li&gt;
&lt;li&gt;5,000 different queries per run&lt;/li&gt;
&lt;li&gt;A working set that fits in available memory&lt;/li&gt;
&lt;li&gt;The default oversampling and rescoring behavior&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The initial comparison used row-wise and columnar vector storage:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmnt.cr%2Fgo%2F3CqVZn" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmnt.cr%2Fgo%2F3CqVZn" alt="KNN throughput with row-wise and columnar vector storage" width="700" height="410"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Across the three measurements, the row-wise path delivered 2.80x - 3.53x the throughput.&lt;/p&gt;

&lt;p&gt;Graph traversal contributes equally in both results. The throughput gap comes from the rescoring pass.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the defaults make this important
&lt;/h2&gt;

&lt;p&gt;Rescoring and oversampling are part of Manticore's default KNN behavior:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;oversampling=3.0&lt;/code&gt; multiplies the requested &lt;code&gt;k&lt;/code&gt; before HNSW search, retrieving more approximate candidates than the final query needs.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;rescore=1&lt;/code&gt; fetches the original 32-bit vectors for those candidates, recalculates their distances, sorts them again, and returns the final top &lt;code&gt;k&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The default query path is therefore:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;requested k -&amp;gt; retrieve up to 3 x k candidates -&amp;gt; rescoring -&amp;gt; return k results
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That changes how the benchmark values should be interpreted:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Requested k&lt;/th&gt;
&lt;th&gt;Target HNSW candidate pool&lt;/th&gt;
&lt;th&gt;Final results after rescoring&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;60&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;300&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;500&lt;/td&gt;
&lt;td&gt;1,500&lt;/td&gt;
&lt;td&gt;500&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For example, &lt;code&gt;k=500&lt;/code&gt; makes HNSW search with an effective &lt;code&gt;k&lt;/code&gt; of 1,500. Those candidates become eligible for exact rescoring, and the best 500 are returned. Filtering, disk chunks, and candidate availability can affect the exact number of physical reads, while the target candidate pool remains three times the requested result count.&lt;/p&gt;

&lt;p&gt;Oversampling and rescoring improve ranking quality, particularly with quantized vectors. Disabling them changes the normal quality/performance tradeoff. Optimizing the default path benefits typical KNN queries directly.&lt;/p&gt;

&lt;h2&gt;
  
  
  The graph search is the same; the vector access is different
&lt;/h2&gt;

&lt;p&gt;Row-wise and columnar KNN tables use the same HNSW graph search. In both scenarios, the graph navigates the approximate index and produces candidate document IDs. Vector storage becomes relevant after that stage, when rescoring needs the original full-precision value for every candidate.&lt;/p&gt;

&lt;p&gt;The row-wise accessor can provide a stable address for each resident vector. Manticore can retain several vector pointers, prefetch their data, and calculate multiple distances together.&lt;/p&gt;

&lt;p&gt;Columnar access works differently. It fetches a vector through a reusable read buffer. A later read can overwrite that buffer, so the rescoring code processes one vector before fetching the next instead of retaining a group of pointers for batch processing.&lt;/p&gt;

&lt;p&gt;This distinction grows more expensive with high-dimensional vectors and as the user increases &lt;code&gt;k&lt;/code&gt;. The benchmark illustrates the effect at effective candidate pools of 60, 300, and 1,500, but users can request other &lt;code&gt;k&lt;/code&gt; values and the same mechanism applies.&lt;/p&gt;

&lt;p&gt;Since HNSW traversal is the same, this rescoring behavior accounts for the observed storage-mode gap.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why columnar storage still matters
&lt;/h2&gt;

&lt;p&gt;Columnar storage was originally intended for cases where there is not enough memory to load all the data for a queried attribute. It arranges all values of one attribute next to one another. A page read for &lt;code&gt;price&lt;/code&gt;, for example, contains mostly more &lt;code&gt;price&lt;/code&gt; values rather than prices interleaved with categories, timestamps, and other attributes.&lt;/p&gt;

&lt;p&gt;Row-wise storage keeps the attributes for one document together. That is useful when a query needs the complete row, while a query that reads one attribute may also bring unrelated values into memory. For scans, filters, and aggregations over individual attributes, this lowers the useful-data density of each loaded page.&lt;/p&gt;

&lt;p&gt;The contiguous columnar layout normally behaves better under memory pressure because it reads and caches more of the requested attribute and less unrelated data. That advantage becomes especially important when the table is much larger than RAM.&lt;/p&gt;

&lt;p&gt;The slowdown had a simpler cause: random vector reads during rescoring had to pass through one reusable buffer.&lt;/p&gt;

&lt;h2&gt;
  
  
  A new access path for columnar vectors
&lt;/h2&gt;

&lt;p&gt;Memory mapping provides stable access to the columnar files.&lt;/p&gt;

&lt;p&gt;Mapping reserves virtual address space for a file. Physical pages enter RAM on demand as the process accesses them, and the operating system can reclaim those pages under memory pressure. A mapping can therefore be larger than the available physical memory.&lt;/p&gt;

&lt;p&gt;For rescoring, pointer stability is the key change. A vector can be addressed directly in the mapped region instead of copied through a reusable read buffer. Manticore can then:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Sort candidates by disk chunk and row ID to improve locality.&lt;/li&gt;
&lt;li&gt;Collect stable vector pointers in batches of up to 256 candidates.&lt;/li&gt;
&lt;li&gt;Prefetch vector data before it is needed.&lt;/li&gt;
&lt;li&gt;Calculate distances for the batch, processing pairs together where supported and handling any remainder individually.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Memory mapping enables the batched rescoring path that row-wise storage already benefited from.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fit-in-memory result: most of the gap disappears
&lt;/h2&gt;

&lt;p&gt;We repeated the DBpedia test with memory-mapped columnar access and compared all three storage paths:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmnt.cr%2Fgo%2FapKqxJ" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmnt.cr%2Fgo%2FapKqxJ" alt="KNN throughput by vector storage mode" width="700" height="430"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;At &lt;code&gt;k=20&lt;/code&gt;, columnar throughput rose from 178 to 509 QPS, or 2.86 times the &lt;code&gt;file&lt;/code&gt; result. At &lt;code&gt;k=100&lt;/code&gt;, it rose from 97 to 304 QPS, a 3.13-times improvement. At &lt;code&gt;k=500&lt;/code&gt;, it rose from 61 to 157 QPS, a 2.57-times improvement.&lt;/p&gt;

&lt;p&gt;Columnar &lt;code&gt;file&lt;/code&gt; access reached 28-36% of row-wise throughput. Memory-mapped columnar access reached 85-92%. Its remaining gap to row-wise storage was 15.4% at &lt;code&gt;k=20&lt;/code&gt;, 11.1% at &lt;code&gt;k=100&lt;/code&gt;, and 8.2% at &lt;code&gt;k=500&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The narrowing gap is consistent with batching becoming more valuable as &lt;code&gt;k&lt;/code&gt; and the resulting rescoring work increase. At the benchmark's &lt;code&gt;k=500&lt;/code&gt; point, memory-mapped columnar storage finished within about 8% of row-wise performance instead of delivering roughly one-third of its throughput.&lt;/p&gt;

&lt;p&gt;That establishes the result for a resident working set. The next question is how the new path behaves in the memory-constrained conditions columnar storage was designed for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Out-of-core result: the taxi benchmark
&lt;/h2&gt;

&lt;p&gt;For generic search, we used a much larger taxi dataset on the same Ryzen 9 5950X:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;1.74 billion documents&lt;/li&gt;
&lt;li&gt;32 disk chunks&lt;/li&gt;
&lt;li&gt;372 GB total table size&lt;/li&gt;
&lt;li&gt;88 GB of queried &lt;code&gt;.spc&lt;/code&gt; columnar storage files&lt;/li&gt;
&lt;li&gt;32 GB Docker memory limit&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The queried columnar footprint was 2.75 times the container's entire memory limit. The search server also needed memory for its own data structures, leaving less than 32 GB for filesystem-backed pages. This is an out-of-core workload, so queries run under page eviction and physical I/O.&lt;/p&gt;

&lt;p&gt;We ran the same generic-search suite against two versions of the data:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;taxi:&lt;/strong&gt; the complete 1.74-billion-document table, which exceeds available memory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;taxi1:&lt;/strong&gt; one disk chunk, whose working set fits in memory.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The 17 queries included full-text search, unfiltered aggregates, equality and range filters, indexed lookups, and high- and low-cardinality &lt;code&gt;GROUP BY&lt;/code&gt; operations. These are the same taxi queries used in our public comparisons on &lt;a href="https://mnt.cr/go/9aIgaA" rel="noopener noreferrer"&gt;db-benchmarks.com&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Each access path was tested in three complete runs. The figures below use the arithmetic mean of server-reported time across those runs, with minimum and maximum values shown by the whiskers. Each value represents the total for the entire 17-query suite.&lt;/p&gt;

&lt;p&gt;Cold and hot measurements are analyzed separately:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cold&lt;/strong&gt; is the first measured execution in the benchmark's cache-drop phase.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hot&lt;/strong&gt; is the mean of 10 repeated executions per query on taxi and 50 on taxi1.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmnt.cr%2Fgo%2FtSfKR1" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmnt.cr%2Fgo%2FtSfKR1" alt="Cold and hot generic-search time with columnar file and mmap access" width="700" height="430"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Cold runs
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dataset&lt;/th&gt;
&lt;th&gt;Columnar access&lt;/th&gt;
&lt;th&gt;Average&lt;/th&gt;
&lt;th&gt;Min-max&lt;/th&gt;
&lt;th&gt;Change vs file&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Whole taxi table, out of core&lt;/td&gt;
&lt;td&gt;&lt;code&gt;file&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;25.086 s&lt;/td&gt;
&lt;td&gt;24.489-25.581 s&lt;/td&gt;
&lt;td&gt;-&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Whole taxi table, out of core&lt;/td&gt;
&lt;td&gt;&lt;code&gt;mmap&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;25.155 s&lt;/td&gt;
&lt;td&gt;24.561-25.952 s&lt;/td&gt;
&lt;td&gt;+0.28%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Single taxi chunk, resident&lt;/td&gt;
&lt;td&gt;&lt;code&gt;file&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;845.667 ms&lt;/td&gt;
&lt;td&gt;793-905 ms&lt;/td&gt;
&lt;td&gt;-&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Single taxi chunk, resident&lt;/td&gt;
&lt;td&gt;&lt;code&gt;mmap&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;832.667 ms&lt;/td&gt;
&lt;td&gt;815-848 ms&lt;/td&gt;
&lt;td&gt;-1.54%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Lower is better.&lt;/p&gt;

&lt;p&gt;On the complete taxi table, mmap was 0.28% slower in cold runs, a negligible difference.&lt;/p&gt;

&lt;p&gt;On the resident single chunk, mmap was 1.54% faster: 832.667 milliseconds versus 845.667 milliseconds, which is a small improvement.&lt;/p&gt;

&lt;h2&gt;
  
  
  Hot runs
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dataset&lt;/th&gt;
&lt;th&gt;Columnar access&lt;/th&gt;
&lt;th&gt;Average&lt;/th&gt;
&lt;th&gt;Min-max&lt;/th&gt;
&lt;th&gt;Change vs file&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Whole taxi table, out of core&lt;/td&gt;
&lt;td&gt;&lt;code&gt;file&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;22.196 s&lt;/td&gt;
&lt;td&gt;22.033-22.521 s&lt;/td&gt;
&lt;td&gt;-&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Whole taxi table, out of core&lt;/td&gt;
&lt;td&gt;&lt;code&gt;mmap&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;22.223 s&lt;/td&gt;
&lt;td&gt;22.074-22.500 s&lt;/td&gt;
&lt;td&gt;+0.12%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Single taxi chunk, resident&lt;/td&gt;
&lt;td&gt;&lt;code&gt;file&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;758.667 ms&lt;/td&gt;
&lt;td&gt;754.120-765.920 ms&lt;/td&gt;
&lt;td&gt;-&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Single taxi chunk, resident&lt;/td&gt;
&lt;td&gt;&lt;code&gt;mmap&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;751.087 ms&lt;/td&gt;
&lt;td&gt;742.660-759.540 ms&lt;/td&gt;
&lt;td&gt;-1.00%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Lower is better.&lt;/p&gt;

&lt;p&gt;On the out-of-core table, mmap was 0.12% slower in hot runs, a negligible difference.&lt;/p&gt;

&lt;p&gt;On the resident chunk, mmap was 1.00% faster: 751.087 milliseconds versus 758.667 milliseconds. Grouping and range queries included improvements, while some simple aggregates moved slightly in the other direction. The mixed per-query changes and small aggregate difference again support near-parity.&lt;/p&gt;

&lt;p&gt;Together, the cold and hot runs show that the mapped path preserves non-KNN search performance. It was neutral when the queried columnar files exceeded available memory and slightly faster when one chunk fit in memory.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why non-KNN search remains close
&lt;/h2&gt;

&lt;p&gt;The large KNN improvement comes from changing the rescoring strategy: stable vector pointers enable better locality, prefetching, and batched distance calculations. The non-KNN taxi queries do not use this batched KNN rescoring path, so they don't benefit from it.&lt;/p&gt;

&lt;p&gt;On Linux, normal reads and file-backed mmap both go through the kernel page cache. The &lt;code&gt;file&lt;/code&gt; mode reads cached data into an application buffer. &lt;code&gt;mmap&lt;/code&gt; exposes the file-backed pages through the process address space and brings missing pages in through page faults. Both paths remain subject to the same physical-memory limit, page reclaim, and backing storage.&lt;/p&gt;

&lt;p&gt;This shared foundation helps explain the non-KNN parity. &lt;code&gt;mmap&lt;/code&gt; changes the pointer and access model enough to unlock batched KNN rescoring, while the operating system continues to manage the underlying file pages under both access modes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Rescoring can be a substantial part of default KNN execution. Three-times oversampling means a query asking for 500 results uses an effective HNSW &lt;code&gt;k&lt;/code&gt; of 1,500 before the rescoring pass. As users raise or lower &lt;code&gt;k&lt;/code&gt;, the rescoring workload changes with it. When candidate vectors were read one at a time through columnar &lt;code&gt;file&lt;/code&gt; access, KNN throughput was 64-72% lower than with row-wise storage, even though HNSW traversal itself was unchanged.&lt;/p&gt;

&lt;p&gt;Stable mapped addresses allow Manticore to batch that work. On DBpedia, columnar throughput rose by 2.57-3.13 times and reached 85-92% of row-wise performance.&lt;/p&gt;

&lt;p&gt;The memory-constrained tests show that the approach also preserves columnar storage's original strength. With 88 GB of queried columnar data under a 32 GB container limit, non-KNN search time changed by only +0.28% cold and +0.12% hot. The resident single-chunk results were slightly favorable at -1.54% cold and -1.00% hot. The overall result is a large improvement to default KNN rescoring with non-KNN search parity under both memory conditions. As a result, &lt;code&gt;mmap&lt;/code&gt; is now the default access mode for columnar storage.&lt;/p&gt;

&lt;h2&gt;
  
  
  Columnar access configuration
&lt;/h2&gt;

&lt;p&gt;The &lt;code&gt;access_columnar_attrs&lt;/code&gt; option controls how columnar files are accessed. Its default value is &lt;code&gt;mmap&lt;/code&gt;, so no per-table setting is required. The mode can also be selected explicitly when creating a table:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;products&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;title&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;embedding&lt;/span&gt; &lt;span class="n"&gt;FLOAT_VECTOR&lt;/span&gt;
        &lt;span class="n"&gt;KNN_TYPE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'hnsw'&lt;/span&gt;
        &lt;span class="n"&gt;KNN_DIMS&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'1536'&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;ENGINE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'columnar'&lt;/span&gt;
  &lt;span class="n"&gt;access_columnar_attrs&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'mmap'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The previous &lt;code&gt;file&lt;/code&gt; mode remains available. To use it as the server-wide default, add it to the &lt;code&gt;searchd&lt;/code&gt; section of the configuration file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ini"&gt;&lt;code&gt;&lt;span class="err"&gt;searchd&lt;/span&gt; &lt;span class="err"&gt;{&lt;/span&gt;
    &lt;span class="py"&gt;access_columnar_attrs&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;file&lt;/span&gt;
&lt;span class="err"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The option changes how columnar files are accessed. Their stored format and the KNN query syntax stay the same.&lt;/p&gt;

</description>
      <category>database</category>
      <category>performance</category>
      <category>linux</category>
      <category>opensource</category>
    </item>
    <item>
      <title>How search works in Kiva's SaaS, where every customer has their own database</title>
      <dc:creator>Sergey Nikolaev</dc:creator>
      <pubDate>Fri, 02 Oct 2026 12:00:01 +0000</pubDate>
      <link>https://dev.to/sanikolaev/how-search-works-in-kivas-saas-where-every-customer-has-their-own-database-29nf</link>
      <guid>https://dev.to/sanikolaev/how-search-works-in-kivas-saas-where-every-customer-has-their-own-database-29nf</guid>
      <description>&lt;p&gt;Originally published on the &lt;a href="https://mnt.cr/go/KbgfHn" rel="noopener noreferrer"&gt;Manticore Search website&lt;/a&gt; on October 1, 2026&lt;/p&gt;

&lt;h1&gt;
  
  
  How search works in Kiva's SaaS, where every customer has their own database
&lt;/h1&gt;

&lt;p&gt;How Kiva uses Manticore Search in a SaaS platform where every customer has their own database and schema: combining a periodically rebuilt table with an RT table, processing changes in real time, and running search without dedicated servers.&lt;/p&gt;

&lt;p&gt;In a SaaS platform that customers configure with low-code/no-code tools, it is difficult to know in advance exactly what data users will want to search. One customer adds their own fields and forms, another configures custom business processes, and a third uses the data they find for filtering, bulk editing, or marketing campaigns.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://mnt.cr/go/cKx59q" rel="noopener noreferrer"&gt;Kiva Teknoloji&lt;/a&gt; faced exactly this challenge. The Turkish company has been developing cloud business applications since 2009. Its main product, KivaCRM, is built on the company's own Kiva Cloud Platform and combines CRM with tools for business process automation, reporting, and analytics. Customers can configure forms, lists, processes, reports, and dashboards themselves, so the same platform is used by companies across more than 40 industries.&lt;/p&gt;

&lt;p&gt;That flexibility directly affects search. Every Kiva customer has their own database, and the table structure changes as they configure the application around their own processes. This means search cannot be configured once for a fixed set of fields and then left unchanged.&lt;/p&gt;

&lt;p&gt;Manticore Search works here as a separate search layer alongside MySQL. It handles full-text search, while MySQL remains the primary database. This approach has allowed Kiva to reduce the load on its main database servers, add more flexible search capabilities, and do so with almost no increase in infrastructure costs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Search is part of the platform, not just a search box
&lt;/h2&gt;

&lt;p&gt;Kiva indexes all kinds of data that customers create in their applications. Search is available in different parts of the product, and the records it finds are used for more than just viewing.&lt;/p&gt;

&lt;p&gt;According to &lt;a href="https://mnt.cr/go/rS2gtl" rel="noopener noreferrer"&gt;Arda Beyazoglu&lt;/a&gt;, Lead Engineer at Kiva Teknoloji, search is used for analytics, filtering, bulk record editing, email and SMS campaigns, and other tasks.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"We use Manticore for full-text search and index all kinds of data generated by our customers."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is not a typical catalog search with a predefined schema. Every Kiva customer has a separate database, and its structure depends on how the application is configured. As customers customize it, new fields and entities appear.&lt;/p&gt;

&lt;p&gt;So the search layer has to adapt to each customer's data rather than to a single predefined model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why MySQL full-text search was not enough
&lt;/h2&gt;

&lt;p&gt;Before moving to Manticore, Kiva used MySQL's built-in full-text search.&lt;/p&gt;

&lt;p&gt;According to Arda, it was slower than Manticore and consumed resources on the main database server. That matters in a SaaS environment: the application itself needs the same CPU and memory, so search starts competing with its primary workload.&lt;/p&gt;

&lt;p&gt;Kiva therefore moved &lt;a href="https://mnt.cr/go/FM4CNB" rel="noopener noreferrer"&gt;full-text search&lt;/a&gt; into a separate layer instead of making MySQL handle both the main application workload and search.&lt;/p&gt;

&lt;p&gt;MySQL remained the primary data store, while Manticore took over search.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Kiva moved to Manticore
&lt;/h2&gt;

&lt;p&gt;Before Manticore, the team used Sphinx for some time. Kiva later moved to Manticore because of version compatibility issues and the slower pace of development of the previous solution. The team was still able to keep its existing search architecture.&lt;/p&gt;

&lt;h2&gt;
  
  
  A separate search table for every customer
&lt;/h2&gt;

&lt;p&gt;The architecture now looks like this:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;customer MySQL → search cache → periodically rebuilt table + RT table → Manticore distributed table → search → record IDs → MySQL&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;For every table containing customer data, Kiva generates a search cache. It then creates a separate &lt;a href="https://mnt.cr/go/keeBPm" rel="noopener noreferrer"&gt;Manticore distributed table&lt;/a&gt; for each customer.&lt;/p&gt;

&lt;p&gt;It combines two local tables:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;one is fully rebuilt once a week;&lt;/li&gt;
&lt;li&gt;real-time changes go into an &lt;a href="https://mnt.cr/go/xFycRU" rel="noopener noreferrer"&gt;RT table&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This lets users search up-to-date data without Kiva having to constantly rebuild the entire search table.&lt;/p&gt;

&lt;p&gt;When a user searches in the application, Manticore finds the matching records. The application then performs point lookups in the source MySQL tables to retrieve the rest of the data.&lt;/p&gt;

&lt;p&gt;This avoids duplicating the full contents of the source tables in Manticore. Search only needs to return the IDs of matching records, while the application retrieves the rest of the data from MySQL.&lt;/p&gt;

&lt;p&gt;Each system therefore has a clear role:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;MySQL stores the application's source data;&lt;/li&gt;
&lt;li&gt;Manticore handles search;&lt;/li&gt;
&lt;li&gt;the application uses IDs to connect search results back to the source records.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  No dedicated search servers required
&lt;/h2&gt;

&lt;p&gt;The amount of data varies significantly between Kiva customers: the platform is used by both small businesses and large enterprises.&lt;/p&gt;

&lt;p&gt;Arda gives the following approximate figures:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Size of most tables&lt;/td&gt;
&lt;td&gt;up to a few GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A few large tables&lt;/td&gt;
&lt;td&gt;30 GB and above&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Full rebuild of smaller tables&lt;/td&gt;
&lt;td&gt;a few minutes at most&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Full rebuild of larger tables&lt;/td&gt;
&lt;td&gt;about 30–60 minutes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dedicated server for Manticore&lt;/td&gt;
&lt;td&gt;not used by Kiva&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The last point is especially notable.&lt;/p&gt;

&lt;p&gt;Kiva does not provision dedicated servers for Manticore. The search engine runs either on an application server or on a server hosting a database replica.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"We never run Manticore on a separate server. It uses very few resources when idle and is very CPU-efficient, so at our scale the additional infrastructure cost is almost zero."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For Kiva, this means that adding a separate search layer did not turn into another cluster that has to be paid for and maintained.&lt;/p&gt;

&lt;h2&gt;
  
  
  Less load on MySQL, more resources for the application
&lt;/h2&gt;

&lt;p&gt;The main benefit for Kiva is not just faster full-text search.&lt;/p&gt;

&lt;p&gt;Moving the search workload out of MySQL freed up resources on the database servers. Those resources can be used for more important application data and operations, while the same infrastructure can now serve more customers.&lt;/p&gt;

&lt;p&gt;Manticore also gave Kiva capabilities it did not have before, including more advanced search methods and support for more languages.&lt;/p&gt;

&lt;h2&gt;
  
  
  Next step: hybrid search
&lt;/h2&gt;

&lt;p&gt;Kiva is also considering Manticore for new applications with AI features.&lt;/p&gt;

&lt;p&gt;According to Arda, the team is considering &lt;a href="https://mnt.cr/go/Kt2dkO" rel="noopener noreferrer"&gt;hybrid search&lt;/a&gt;, which combines full-text and vector search. For Kiva, this is a natural extension of the existing architecture: Manticore already serves as the search layer for customer data, so new search methods can be added there without moving the source data out of MySQL.&lt;/p&gt;

&lt;p&gt;For now, this is only a plan, not something already running in production. But it shows how the search layer can evolve from full-text search toward search that considers both the words in a query and their meaning.&lt;/p&gt;

&lt;h2&gt;
  
  
  Searching data with a changing structure
&lt;/h2&gt;

&lt;p&gt;At Kiva, Manticore fits into the existing infrastructure and does not require a separate search cluster.&lt;/p&gt;

&lt;p&gt;Every customer has their own database and a schema that changes as the application is configured. Manticore combines a periodically rebuilt table with changes from the RT table and returns the IDs of matching records. MySQL remains the primary data store.&lt;/p&gt;

&lt;p&gt;As a result, Kiva has moved the full-text search workload off its primary database, made more efficient use of its existing servers, and turned search into a shared platform capability — even when it is impossible to know in advance what fields and entities the next customer will create.&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>database</category>
      <category>mysql</category>
      <category>search</category>
    </item>
    <item>
      <title>How wine.co.za made an archive of about 145,000 records easier to search with Manticore Search</title>
      <dc:creator>Sergey Nikolaev</dc:creator>
      <pubDate>Tue, 29 Sep 2026 13:00:04 +0000</pubDate>
      <link>https://dev.to/sanikolaev/how-winecoza-made-an-archive-of-about-145000-records-easier-to-search-with-manticore-search-5c1e</link>
      <guid>https://dev.to/sanikolaev/how-winecoza-made-an-archive-of-about-145000-records-easier-to-search-with-manticore-search-5c1e</guid>
      <description>&lt;p&gt;Originally published on the &lt;a href="https://mnt.cr/go/8dogb4" rel="noopener noreferrer"&gt;Manticore Search website&lt;/a&gt; on September 29, 2026.&lt;/p&gt;

&lt;h1&gt;
  
  
  How wine.co.za made an archive of about 145,000 records easier to search with Manticore Search
&lt;/h1&gt;

&lt;p&gt;How wine.co.za uses Manticore Search alongside MariaDB to improve search across its wine industry archive, provide better suggestions, and keep maintenance to a minimum.&lt;/p&gt;

&lt;p&gt;For &lt;a href="https://mnt.cr/go/xM7OyP" rel="noopener noreferrer"&gt;wine.co.za&lt;/a&gt;, search is about more than finding a bottle of wine.&lt;/p&gt;

&lt;p&gt;The site's archive covers much of the South African wine industry: wines, wineries, contacts, international agents, news articles, and classified adverts. Visitors might be looking for a particular producer, an old article, or a wine they can describe but don't know by name.&lt;/p&gt;

&lt;p&gt;wine.co.za used to handle these searches directly in MariaDB, using full-text indexes and wildcard queries. It worked, but visitors had to be fairly precise about what they typed.&lt;/p&gt;

&lt;p&gt;Manticore Search made those searches more flexible and made search suggestions more useful. For &lt;a href="https://mnt.cr/go/Qeu9RQ" rel="noopener noreferrer"&gt;Kevin Kidson&lt;/a&gt;, IT Director at WineNet, the other big benefit is that it needs very little attention:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“It is just so damn solid — build and forget — it has been running rock solid for years.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  An archive of the South African wine industry
&lt;/h2&gt;

&lt;p&gt;Manticore powers search across all parts of wine.co.za. When Kevin shared details of the system, it indexed the following:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dataset&lt;/th&gt;
&lt;th&gt;Records&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Wineries&lt;/td&gt;
&lt;td&gt;907&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Wines&lt;/td&gt;
&lt;td&gt;45,833&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Contacts&lt;/td&gt;
&lt;td&gt;4,644&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;International agents&lt;/td&gt;
&lt;td&gt;3,447&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;News articles&lt;/td&gt;
&lt;td&gt;41,602&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Classified adverts&lt;/td&gt;
&lt;td&gt;48,482&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total in these collections&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;144,915&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;There are a few smaller databases as well. Kevin describes wine.co.za as an archive for the South African wine industry, where every record matters.&lt;/p&gt;

&lt;p&gt;The aim was straightforward: make all this information easier to find with a search engine that was easy to set up and could work alongside the existing application.&lt;/p&gt;

&lt;h2&gt;
  
  
  Before Manticore: visitors had to know what to type
&lt;/h2&gt;

&lt;p&gt;Kevin describes the old search as a combination of full-text indexes and wildcard queries using patterns such as &lt;code&gt;xx*&lt;/code&gt; and &lt;code&gt;xx%&lt;/code&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“The user had to know what they were looking for, or at least how to start spelling it.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That was limiting for someone who knew what kind of wine they wanted but didn't have a particular name in mind. They might want to search for:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;dry chenin stellenbosch&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Those three words describe a wine style, a grape variety, and a region. Kevin says the ability to combine them in one search was the most important improvement Manticore brought to the site.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“It has allowed me to combine search words, so dry chenin from Stellenbosch etc — I could never do that before.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Visitors can now describe what they want rather than having to start with a specific wine or winery name.&lt;/p&gt;

&lt;h2&gt;
  
  
  MariaDB remains the source of truth
&lt;/h2&gt;

&lt;p&gt;The integration is simple. wine.co.za still stores and updates its data in MariaDB. Once a day, it copies the searchable data into Manticore. The website sends search queries through the .NET library.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MariaDB → daily indexing → Manticore Search → .NET application → wine.co.za search&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;MariaDB continues to handle the primary data, while Manticore handles search. The two work alongside each other rather than one replacing the other.&lt;/p&gt;

&lt;p&gt;Ease of setup mattered because Kevin had little time to spare:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“It was very easy and logical to set up.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;He remembers the setup being spread over about two days because he was fitting it around other work. He estimates the total time spent installing Manticore and learning how to use it at about 24 hours.&lt;/p&gt;

&lt;h2&gt;
  
  
  Better suggestions as visitors type
&lt;/h2&gt;

&lt;p&gt;Manticore also powers suggestions in wine.co.za's search boxes. These help visitors find what they're looking for before they've finished typing.&lt;/p&gt;

&lt;p&gt;Kevin says Manticore is fast enough to generate suggestions as people type, and its more forgiving search makes the dropdowns easier to build. Visitors aren't tied to a particular word order: they can type &lt;code&gt;dry chenin stellenbosch&lt;/code&gt; or &lt;code&gt;stellenbosch chenin dry&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That helps people who know the exact name of a wine or producer, as well as those who only know a few details about what they want. Both can use the same search box.&lt;/p&gt;

&lt;p&gt;In his survey response, Kevin highlighted both improvements:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“MUCH better searching across the site — and MUCH better suggestions in drop-downs etc.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Years of reliable search
&lt;/h2&gt;

&lt;p&gt;wine.co.za hasn't shared detailed performance benchmarks. Kevin says the company is small and doesn't measure much, so there are no before-and-after latency or resource-use figures to report.&lt;/p&gt;

&lt;p&gt;For Kevin, the more important point is that Manticore has been running reliably for years without needing much of his time. He would like to explore more of what it can do, but he's already happy with how it works.&lt;/p&gt;

&lt;p&gt;For a small team, that's a practical benefit. Time that might otherwise go into maintaining search can go into other work.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“It's doing a brilliant job &amp;amp; I don't have to worry about it.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Better search without a big engineering project
&lt;/h2&gt;

&lt;p&gt;wine.co.za's setup shows how a separate search engine can fit into an existing application without replacing the main database. MariaDB still stores the data. Manticore receives a copy each day, and the .NET application uses it for search.&lt;/p&gt;

&lt;p&gt;The improvement is visible in ordinary searches. Someone looking for a dry Chenin from Stellenbosch can put those words together, try a different order, and get suggestions as they type. They don't have to know a specific wine name before they begin.&lt;/p&gt;

&lt;p&gt;For Kevin, the value is in that combination: visitors have an easier time finding what they need, while search requires very little day-to-day attention from him.&lt;/p&gt;

</description>
      <category>database</category>
      <category>search</category>
      <category>mariadb</category>
      <category>dotnet</category>
    </item>
    <item>
      <title>How Bookbot powers 3.5 million saved searches and ~10% of revenue with Manticore Search</title>
      <dc:creator>Sergey Nikolaev</dc:creator>
      <pubDate>Thu, 17 Sep 2026 09:00:00 +0000</pubDate>
      <link>https://dev.to/sanikolaev/how-bookbot-powers-35-million-saved-searches-and-10-of-revenue-with-manticore-search-4d41</link>
      <guid>https://dev.to/sanikolaev/how-bookbot-powers-35-million-saved-searches-and-10-of-revenue-with-manticore-search-4d41</guid>
      <description>&lt;p&gt;Originally published on &lt;a href="https://mnt.cr/go/wx99Cf" rel="noopener noreferrer"&gt;https://mnt.cr/go/wx99Cf&lt;/a&gt; on September 15, 2026&lt;/p&gt;

&lt;h1&gt;
  
  
  How Bookbot powers 3.5 million saved searches and ~10% of revenue with Manticore Search
&lt;/h1&gt;

&lt;p&gt;How Bookbot uses Manticore Search to keep 1.29 million second-hand books searchable, match new arrivals against 3.5 million saved searches, and power a channel responsible for ~10% of revenue.&lt;/p&gt;

&lt;p&gt;For a second-hand bookstore, search has to do more than find a title. It has to find a copy that actually exists.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://mnt.cr/go/t7XTho" rel="noopener noreferrer"&gt;Bookbot&lt;/a&gt;, known as Knihobot in the Czech Republic, takes in used books, photographs them, catalogs them, sells them, and ships them to their next readers. Two copies of the same title may have different editions, prices, conditions, and photographs. When one copy sells, the search result may need to change immediately-even if the title itself remains available.&lt;/p&gt;

&lt;p&gt;Bookbot uses Manticore Search to keep this constantly changing inventory searchable. Its main book index contains around &lt;strong&gt;5.3 million documents&lt;/strong&gt;, including approximately &lt;strong&gt;1.29 million physical copies currently in stock&lt;/strong&gt;. Manticore also stores around &lt;strong&gt;3.5 million saved-search queries&lt;/strong&gt; for Bookbot's &lt;a href="https://mnt.cr/go/HzTZsK" rel="noopener noreferrer"&gt;Watchdog&lt;/a&gt; feature.&lt;/p&gt;

&lt;p&gt;According to Bookbot's internal figures, Watchdogs account for roughly &lt;strong&gt;10% of revenue&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Bookbot co-founder and CTO &lt;a href="https://mnt.cr/go/kSr4O9" rel="noopener noreferrer"&gt;David Gazdoš&lt;/a&gt; shared how the company uses Manticore not only to power customer search, but also to keep search results aligned with physical inventory, choose the right copy to show, match newly available books with waiting readers, search a 203-million-record bibliographic dataset, and support internal operations.&lt;/p&gt;

&lt;h2&gt;
  
  
  When every copy is different
&lt;/h2&gt;

&lt;p&gt;A conventional online store can often represent a product with one record and a stock counter.&lt;/p&gt;

&lt;p&gt;Bookbot's inventory is different. A title can have several editions, and each edition can have many individual physical copies. Those copies can differ in price and condition, and each has its own photograph and availability.&lt;/p&gt;

&lt;p&gt;That means a Bookbot search result cannot simply say, "we have this title." It needs to represent an offer the customer can actually buy.&lt;/p&gt;

&lt;p&gt;Imagine that the copy currently shown for a title costs €6. That copy sells, but another copy of the same title is still available for €9.&lt;/p&gt;

&lt;p&gt;The title can remain in the results, but the offer needs to change: new price, new photograph, and potentially a different edition. If the customer filtered for books below €8, the title should disappear from that result set altogether.&lt;/p&gt;

&lt;p&gt;This is where Manticore's &lt;a href="https://mnt.cr/go/ZQifBv" rel="noopener noreferrer"&gt;real-time tables&lt;/a&gt; are important.&lt;/p&gt;

&lt;p&gt;MySQL remains the source of truth for Bookbot's business data. Kafka carries events generated by orders, reservations, price changes, lost or discarded books, temporary withdrawals, and catalogue corrections. As the application processes those events, it can &lt;a href="https://mnt.cr/go/CY3TLx" rel="noopener noreferrer"&gt;replace&lt;/a&gt;, update, or remove individual records in Manticore.&lt;/p&gt;

&lt;p&gt;Instead of treating search as a periodically rebuilt copy of the catalogue, Bookbot can keep the searchable offer aligned with what is happening to the physical books.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Manticore lets us update the searchable offer as the physical inventory changes. When one copy sells, we can replace or remove that record without rebuilding the whole index.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For the customer, the benefit is simple: the price, edition, condition, photograph, and availability shown by search continue to describe something they can actually buy.&lt;/p&gt;

&lt;h2&gt;
  
  
  One title in the results, many physical copies underneath
&lt;/h2&gt;

&lt;p&gt;Keeping the inventory fresh solves only half of the problem.&lt;/p&gt;

&lt;p&gt;Customers normally want to browse titles, not dozens of nearly identical copies of the same book. But Bookbot still needs to preserve the differences between those copies underneath.&lt;/p&gt;

&lt;p&gt;Manticore's &lt;a href="https://mnt.cr/go/P4Z5AO" rel="noopener noreferrer"&gt;grouping&lt;/a&gt; lets Bookbot bridge those two levels.&lt;/p&gt;

&lt;p&gt;The application first filters the physical copies that satisfy the customer's request. It then groups the remaining records by title and uses &lt;code&gt;WITHIN GROUP ORDER BY&lt;/code&gt; to choose which copy should represent that title in the result set.&lt;/p&gt;

&lt;p&gt;The order matters.&lt;/p&gt;

&lt;p&gt;Suppose a title has copies from different publishers at €6, €9, and €14. A customer asks for Publisher B, a price below €10, and no recorded damage. The result needs to be a physical copy that satisfies &lt;strong&gt;all three conditions at once&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Bookbot cannot take the publisher from one copy, the price from another, and the condition or photograph from a third. That would describe an offer that does not exist.&lt;/p&gt;

&lt;p&gt;Manticore lets the filters operate on the underlying copy-level records before the matches are grouped. The representative copy can then be selected according to the requested ordering. Sorting by price may pick one copy; sorting by publication year may pick a different edition.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Manticore gives us the copy-level precision we need while still letting us present one clean result per title. The filters apply to the real physical copies first, so the result always represents an offer we can actually sell.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;When that representative copy sells, another suitable copy can take its place.&lt;/p&gt;

&lt;p&gt;The same principle applies to facets. Bookbot counts distinct titles rather than raw physical copies, so a title with hundreds of copies does not dominate a category simply because more of it happens to be in stock.&lt;/p&gt;

&lt;p&gt;For browsing requests that can safely be answered at title level, Bookbot maintains a smaller precomputed index with about &lt;strong&gt;3.59 million title records&lt;/strong&gt;. When a query requires copy-level correctness-such as full-text search or filters that must all match the same physical item-the application uses the full book index instead.&lt;/p&gt;

&lt;p&gt;This gives Bookbot both sides of the tradeoff: a faster path where title-level data is sufficient, and copy-level precision where it is not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Turning unavailable books into future demand
&lt;/h2&gt;

&lt;p&gt;The most interesting use of Manticore at Bookbot begins when there is nothing to sell.&lt;/p&gt;

&lt;p&gt;Second-hand inventory is unpredictable. A reader may know exactly which book or edition they want while Bookbot has no suitable copy available.&lt;/p&gt;

&lt;p&gt;Instead of making that reader come back and repeat the search every few days, Bookbot lets them create a Watchdog. It can track a title, author, genre, edition, price, or a more detailed combination of criteria and notify the reader when a matching book appears.&lt;/p&gt;

&lt;p&gt;This reverses the normal direction of search.&lt;/p&gt;

&lt;p&gt;With ordinary search:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;one query → matching books&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;With a Watchdog:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;one newly available book → matching saved queries&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Manticore's &lt;a href="https://mnt.cr/go/CajFwy" rel="noopener noreferrer"&gt;percolate tables&lt;/a&gt; are built for this "search in reverse" pattern. Instead of storing documents and running a query against them, Bookbot stores the queries and submits a book to find which saved searches it matches.&lt;/p&gt;

&lt;p&gt;At the time of Bookbot's production snapshot, Manticore held around &lt;strong&gt;3.5 million Watchdog rules&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;When a pricing event makes a book available, the application sends its searchable representation for percolate matching. Manticore returns the matching saved searches. Bookbot then applies its own business logic: whether a Watchdog is still active, whether the customer is eligible for a notification, and whether the match should be sent immediately or included in a digest.&lt;/p&gt;

&lt;p&gt;Manticore handles the matching; Bookbot controls the customer experience around it.&lt;/p&gt;

&lt;p&gt;This turns search into something more valuable than a way to browse current inventory. It becomes a way to preserve demand when supply is missing and reconnect that demand with a future copy.&lt;/p&gt;

&lt;p&gt;According to Bookbot, Watchdogs account for around &lt;strong&gt;10% of revenue&lt;/strong&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Manticore lets us match each newly available book against around 3.5 million saved searches. Watchdogs now account for roughly 10% of our revenue, so this is not just a search feature-it has become an important business channel for us.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For Bookbot, this is a direct business result of being able to search in both directions: readers can search the catalogue, and newly available books can search for readers who already want them.&lt;/p&gt;

&lt;p&gt;The same pattern can apply far beyond books-to back-in-stock alerts, classified listings, jobs, property, tickets, price alerts, or any product where new supply needs to be matched with previously expressed intent.&lt;/p&gt;

&lt;h2&gt;
  
  
  Search that helps readers even when they do not type perfectly
&lt;/h2&gt;

&lt;p&gt;Book search also has a relevance problem: people do not always remember titles, author names, translations, or series exactly as they appear in the catalogue.&lt;/p&gt;

&lt;p&gt;Bookbot uses a dedicated Manticore dictionary with roughly &lt;strong&gt;4.52 million edition records&lt;/strong&gt; for typo correction. It includes titles, authors, series, aliases, curated misspellings, localized fields, and German transliteration variants, while excluding identifiers such as ISBNs from correction candidates.&lt;/p&gt;

&lt;p&gt;The dictionary proposes possible editions, but it does not decide what Bookbot can sell. The main inventory index still determines which physical copies are currently available and satisfy the user's filters.&lt;/p&gt;

&lt;p&gt;This separation is useful because the two datasets have different jobs and different freshness requirements. The typo-correction dictionary can be refreshed on its own schedule, while the main inventory index follows live availability.&lt;/p&gt;

&lt;p&gt;For readers, the result is more forgiving search without sacrificing the correctness of the actual offer.&lt;/p&gt;

&lt;h2&gt;
  
  
  One search engine across the life of a book
&lt;/h2&gt;

&lt;p&gt;Customer search and Watchdogs are the clearest user-facing examples, but Bookbot uses Manticore in several other parts of the business.&lt;/p&gt;

&lt;p&gt;Before a physical book can appear in the store, Bookbot needs to identify and catalogue it. A separate bibliographic metadata pool contains approximately &lt;strong&gt;203.44 million reference records&lt;/strong&gt; searchable by ISBN, title, and author.&lt;/p&gt;

&lt;p&gt;Text such as titles, authors, and annotations remains full-text searchable, while structured metadata such as ISBN, publisher, publication year, language, and category IDs uses Manticore's &lt;a href="https://mnt.cr/go/ksHg7U" rel="noopener noreferrer"&gt;columnar storage&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Internal search also runs on Manticore. Bookbot's backoffice indexes cover books, editions, authors, customers, orders, and accounting records; the internal book index used by pricing and lookup workflows contains about &lt;strong&gt;24.7 million records&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Search extends into delivery as well. Another index contains approximately &lt;strong&gt;197,000 pickup points&lt;/strong&gt;, combining text lookup with geographic filtering and distance ordering using Manticore's &lt;a href="https://mnt.cr/go/3eTojM" rel="noopener noreferrer"&gt;geo-search capabilities&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This breadth matters because Bookbot develops its technology in-house. The team can reuse the same search engine for several very different workloads instead of treating customer search, reverse matching, large metadata lookup, internal search, and geo search as completely separate technical problems.&lt;/p&gt;

&lt;p&gt;Manticore becomes part of the path a book follows through the business: identifying it, making it searchable, selecting the right physical copy, matching it with waiting readers, helping employees work with it, and finding a pickup point after it is sold.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“We use Manticore for storefront search, Watchdogs, a 203-million-record bibliographic dataset, internal search, and pickup-point lookup. Reusing the same search technology across these workloads keeps the architecture much simpler for our team.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Production scale with a compact footprint
&lt;/h2&gt;

&lt;p&gt;Bookbot runs four separate Manticore workloads on Kubernetes: customer search, backoffice search, Watchdogs, and the bibliographic metadata pool. They run on separate servers so each workload can be sized independently.&lt;/p&gt;

&lt;p&gt;A production snapshot from September 5, 2026 looked like this:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Production scale&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Main book index&lt;/td&gt;
&lt;td&gt;~5.30 million documents&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;In-stock physical copies&lt;/td&gt;
&lt;td&gt;~1.29 million&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Precomputed title-level index&lt;/td&gt;
&lt;td&gt;~3.59 million records&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Stored Watchdog queries&lt;/td&gt;
&lt;td&gt;~3.50 million&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bibliographic metadata pool&lt;/td&gt;
&lt;td&gt;~203.44 million records&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Instrumented listing operations&lt;/td&gt;
&lt;td&gt;~20.2 million/day&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Equivalent average operation rate&lt;/td&gt;
&lt;td&gt;~234/sec&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Successful item-selection query time&lt;/td&gt;
&lt;td&gt;~7.7 ms mean&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Successful item-selection queries within 50 ms&lt;/td&gt;
&lt;td&gt;98.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Average CPU across four workloads&lt;/td&gt;
&lt;td&gt;~6.6 cores&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Average memory across four workloads&lt;/td&gt;
&lt;td&gt;~40 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The query-volume figure is a count of instrumented Manticore operations, not human searches. A single page request can issue several operations for results, item selection, and facets.&lt;/p&gt;

&lt;p&gt;Likewise, the 7.7 ms measurement covers the application's Manticore call for successful item-selection queries; it is not complete page-load time or end-to-end inventory-update latency.&lt;/p&gt;

&lt;p&gt;Across the four workloads, Bookbot observed no sustained CPU or memory saturation during the measured period. Customer-facing search used most of the CPU, while the 203-million-record metadata workload used less than 0.1 CPU core on average and about 18.6 GiB of memory.&lt;/p&gt;

&lt;p&gt;The important point is not any single benchmark number. It is that Bookbot is using one search technology across live ecommerce search, reverse matching of millions of saved queries, a 203-million-record reference dataset, internal operations, and geo search-while keeping the infrastructure relatively compact.&lt;/p&gt;

&lt;h2&gt;
  
  
  Search that follows the life of a physical book
&lt;/h2&gt;

&lt;p&gt;Bookbot's use of Manticore is unusual because search is closely tied to the lifecycle of each physical item.&lt;/p&gt;

&lt;p&gt;A book arrives. Metadata search helps identify it. A real-time update makes the copy searchable. Filters make sure its price, edition, condition, and photograph stay together. Grouping gives the title one place in the results while selecting a real purchasable copy.&lt;/p&gt;

&lt;p&gt;If no copy is available, a reader can leave a Watchdog behind. When another copy enters the system, Manticore can match that book against millions of saved searches and help Bookbot reconnect new supply with existing demand.&lt;/p&gt;

&lt;p&gt;When the copy sells, search changes with the inventory and another suitable copy can take its place.&lt;/p&gt;

&lt;p&gt;For readers, that means search that reflects what is actually available and remembers what they want when it is not.&lt;/p&gt;

&lt;p&gt;For Bookbot, it means search contributes at several points in the customer journey-and, through Watchdogs, to a channel the company says generates roughly &lt;strong&gt;10% of its revenue&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That is the value Manticore brings here: not simply finding text quickly, but giving Bookbot one search platform for fast-changing inventory, structured filters, grouped results, reverse search, large reference datasets, and geo search - and turning those capabilities into a better buying experience and measurable business value.&lt;/p&gt;

</description>
      <category>database</category>
      <category>search</category>
      <category>architecture</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Better Vector Search for Long Documents: Chunking Inside Manticore Search</title>
      <dc:creator>Sergey Nikolaev</dc:creator>
      <pubDate>Wed, 16 Sep 2026 09:00:05 +0000</pubDate>
      <link>https://dev.to/sanikolaev/better-vector-search-for-long-documents-chunking-inside-manticore-search-4a0m</link>
      <guid>https://dev.to/sanikolaev/better-vector-search-for-long-documents-chunking-inside-manticore-search-4a0m</guid>
      <description>&lt;p&gt;Originally published on &lt;a href="https://mnt.cr/go/cxcaaA" rel="noopener noreferrer"&gt;https://mnt.cr/go/cxcaaA&lt;/a&gt; on September 15, 2026&lt;/p&gt;

&lt;h1&gt;
  
  
  Better Vector Search for Long Documents: Chunking Inside Manticore Search
&lt;/h1&gt;

&lt;p&gt;An embedding model reads only the first few hundred tokens of a document and silently drops the rest. Manticore Search now splits long documents for you at INSERT time: add chunk_strategy to the vector column and pick one of five strategies. No ingest pipeline, no splitter library. On our own manual, recall@5 for deep content went from 55% to 83%.&lt;/p&gt;

&lt;p&gt;Say you are building search over your team's internal documentation — guides, runbooks, postmortems. You have a table with &lt;a href="https://mnt.cr/go/Y10qc6" rel="noopener noreferrer"&gt;auto embeddings&lt;/a&gt;: you insert text, Manticore runs the model and fills the vector column for you. (If that is new to you, start with &lt;a href="https://mnt.cr/go/fKe3uC" rel="noopener noreferrer"&gt;vector search in Manticore&lt;/a&gt;.) You load a 4,000-word document. The insert succeeds. The search works. Everything looks fine.&lt;/p&gt;

&lt;p&gt;Except the model you picked has a 512-token input window, and that document is about 5,000 tokens long. The model read the first 380 words and threw away the other 3,600. Nothing in the document past that point can ever be retrieved, and nothing anywhere told you. The embedding may not represent the document as a whole either.&lt;/p&gt;

&lt;p&gt;Until now, you would usually split the document into several pieces yourself, create embeddings for each one, and then work out how to combine the results if you wanted document search rather than chunk search. Manticore now handles this in the table definition: add &lt;code&gt;chunk_strategy&lt;/code&gt; to the vector column in &lt;code&gt;CREATE TABLE&lt;/code&gt;, and Manticore splits each document into chunks, embeds every chunk, and searches all of them:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;DROP&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;IF&lt;/span&gt; &lt;span class="k"&gt;EXISTS&lt;/span&gt; &lt;span class="n"&gt;docs&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;docs&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;title&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;content&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;chunks&lt;/span&gt; &lt;span class="n"&gt;float_vector_array&lt;/span&gt; &lt;span class="n"&gt;knn_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'hnsw'&lt;/span&gt; &lt;span class="n"&gt;hnsw_similarity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'cosine'&lt;/span&gt;
    &lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'Xenova/all-MiniLM-L6-v2'&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'title,content'&lt;/span&gt;
    &lt;span class="n"&gt;chunk_strategy&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'sentence'&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'256'&lt;/span&gt; &lt;span class="n"&gt;overlap_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'32'&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the whole feature. No ingest pipeline, no splitter library, no second table for chunks, no &lt;code&gt;GROUP BY&lt;/code&gt; to fold chunk hits back into documents.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Five strategies&lt;/strong&gt;: &lt;code&gt;truncate&lt;/code&gt; (the old default), &lt;code&gt;mean&lt;/code&gt;, &lt;code&gt;fixed&lt;/code&gt;, &lt;code&gt;recursive&lt;/code&gt;, &lt;code&gt;sentence&lt;/code&gt;. Set with &lt;code&gt;chunk_strategy&lt;/code&gt; on a model-backed vector column.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;truncate&lt;/code&gt; and &lt;code&gt;mean&lt;/code&gt;&lt;/strong&gt; produce one vector per document and work on a &lt;code&gt;float_vector&lt;/code&gt; column. &lt;strong&gt;&lt;code&gt;fixed&lt;/code&gt;, &lt;code&gt;recursive&lt;/code&gt; and &lt;code&gt;sentence&lt;/code&gt;&lt;/strong&gt; produce many, so they need a &lt;a href="https://mnt.cr/go/LfTF6g" rel="noopener noreferrer"&gt;&lt;code&gt;float_vector_array&lt;/code&gt;&lt;/a&gt; column.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A document is still one search result.&lt;/strong&gt; Chunks compete individually, and Manticore returns the document once, with &lt;code&gt;knn_dist()&lt;/code&gt; reporting the distance to its closest chunk. &lt;code&gt;k&lt;/code&gt; counts documents, not chunks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tuning knobs&lt;/strong&gt;: &lt;code&gt;max_tokens&lt;/code&gt; (chunk size), &lt;code&gt;overlap_tokens&lt;/code&gt; (shared tokens between neighbors), &lt;code&gt;max_chunks&lt;/code&gt; (ceiling per document).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Measured on the Manticore manual&lt;/strong&gt; (189 pages, ~298k words): for content buried past the model's window, recall@5 went from &lt;strong&gt;55.1% → 83.3%&lt;/strong&gt; and MRR from &lt;strong&gt;0.44 → 0.70&lt;/strong&gt;, at ~2.5× the RAM and ~4× the ingest time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Queries are never chunked.&lt;/strong&gt; A query is short enough to embed as a whole; only stored documents are split.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The problem, shown with a small example
&lt;/h2&gt;

&lt;p&gt;Suppose you have four documents:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Backup and restore runbook&lt;/strong&gt; — about 700 words, roughly 900 tokens. Backup schedules, retention, restore drills, credentials, capacity planning. The &lt;em&gt;last&lt;/em&gt; section explains how to rotate the TLS certificate used by the replication port.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Monitoring and alerting guide&lt;/strong&gt; — unrelated.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Getting started with the CLI&lt;/strong&gt; — unrelated.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;TLS and certificates for the HTTP API&lt;/strong&gt; — a short page that is &lt;em&gt;entirely&lt;/em&gt; about certificates, and never mentions rotation or replication.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;You can create the table and add the documents using the commands below.&lt;/p&gt;

&lt;p&gt;So, what we have is: one table, three vector columns with the same source text — one column per strategy. A single &lt;code&gt;INSERT&lt;/code&gt; fills all three, so the comparison conditions are identical:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;DROP&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;IF&lt;/span&gt; &lt;span class="k"&gt;EXISTS&lt;/span&gt; &lt;span class="n"&gt;docs&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;docs&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;title&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;body&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;v_truncate&lt;/span&gt; &lt;span class="n"&gt;float_vector&lt;/span&gt;       &lt;span class="n"&gt;knn_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'hnsw'&lt;/span&gt; &lt;span class="n"&gt;hnsw_similarity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'cosine'&lt;/span&gt;
    &lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'Xenova/all-MiniLM-L6-v2'&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'title,body'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;v_mean&lt;/span&gt;     &lt;span class="n"&gt;float_vector&lt;/span&gt;       &lt;span class="n"&gt;knn_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'hnsw'&lt;/span&gt; &lt;span class="n"&gt;hnsw_similarity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'cosine'&lt;/span&gt;
    &lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'Xenova/all-MiniLM-L6-v2'&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'title,body'&lt;/span&gt; &lt;span class="n"&gt;chunk_strategy&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'mean'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;v_sentence&lt;/span&gt; &lt;span class="n"&gt;float_vector_array&lt;/span&gt; &lt;span class="n"&gt;knn_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'hnsw'&lt;/span&gt; &lt;span class="n"&gt;hnsw_similarity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'cosine'&lt;/span&gt;
    &lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'Xenova/all-MiniLM-L6-v2'&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'title,body'&lt;/span&gt;
    &lt;span class="n"&gt;chunk_strategy&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'sentence'&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'128'&lt;/span&gt; &lt;span class="n"&gt;overlap_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'32'&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Insert the four documents&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;INSERT&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;docs&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;VALUES&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'Backup and restore runbook'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
   &lt;span class="s1"&gt;'Nightly backups run at 02:00 UTC from the standby node. The job snapshots every table directory, writes a manifest, and uploads the result to object storage. Retention is thirty daily copies, twelve monthly copies, and one yearly copy. A restore drill runs on the first Monday of each month against a scratch cluster. The drill counts as passed only when a full-text search over the restored data returns the same document count as production. Anything less is treated as a failed drill and investigated the same week. Before a restore, freeze the target cluster so that no writes land while files are being replaced. Copy the manifest first and verify its checksum. If the checksum does not match, stop: a partial restore is worse than no restore, because the cluster will start and silently serve half the corpus. After the files are in place, unfreeze and let replication catch up. Watch the queue depth. If it does not drain within ten minutes, the node is probably still reading from cold storage and needs a warm-up pass before it can serve traffic. Backup failures page the on-call engineer. The three most common causes are an expired object storage credential, a disk that filled up while the snapshot was being written, and a table left frozen by a previous failed run. All three are recoverable without data loss. Check the job log first, then the disk, then the freeze state of every table. Capacity planning for backups is boring but it matters. A daily copy of the search cluster is roughly the size of the data directory plus fifteen percent for the manifest and metadata. Multiply by the retention count, add the transfer cost, and you have the monthly bill. Most teams discover too late that the yearly copies dominate the storage line. Object storage lifecycle rules do most of the retention work. Daily copies move to infrequent access after seven days and expire after thirty. Monthly copies move to archive after sixty days. Yearly copies never expire automatically; deleting one is a manual action that requires a second approver. Credentials for the backup job live in the secret manager and are issued to a role, not to a person. The role can write new objects and list the bucket. It cannot delete, and it cannot read objects older than the current day. That last restriction is the cheapest defence against a compromised backup runner turning into a data exfiltration path. Documentation for each table lives next to its schema: what the table is for, who owns it, how large it is expected to get, and whether it can be rebuilt from an upstream source. A table that can be rebuilt does not need thirty daily copies. Roughly half of most clusters turns out to be derived data that nobody had marked as derived. Verification is not the same as the job exiting zero. The job can succeed while producing an unusable copy: an empty table, a truncated upload, a manifest that references a file that was never written. The verification step reads the manifest back, checks every referenced object exists and matches its recorded size, and compares row counts on three sampled tables against production. Rotating the replication TLS certificate is a separate procedure and the step people most often get wrong. The certificate that secures the replication port is not the same as the one the HTTP API uses, and replacing one does not replace the other. Generate the new key and signing request on the node that will be rotated first, sign them with the cluster certificate authority, and place the files next to the existing ones rather than on top of them. Then update the node configuration to point at the new paths and reload. Do one node at a time and confirm that the cluster reports every peer as synced before moving on. A half-rotated cluster where two nodes trust different authorities will keep accepting writes on both sides and diverge quietly. When every node has been rotated, remove the old key material and revoke the retired certificate at the authority.'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'Monitoring and alerting guide'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
   &lt;span class="s1"&gt;'Every node exports metrics over an HTTP endpoint that a scraper collects once per fifteen seconds. The dashboards are grouped into four rows: traffic, latency, saturation, and errors. Traffic is queries per second broken down by table. Latency is the ninety-fifth and ninety-ninth percentile of query time, measured server side. Alerting is deliberately thin. Paging alerts fire on sustained error rate above one percent for five minutes, on ninety-ninth percentile latency above two seconds for ten minutes, and on a node dropping out of the cluster. Everything else is a ticket, not a page. Teams that page on every anomaly stop reading pages within a month. Log retention is fourteen days hot and ninety days cold. The query log records the query text, the table, the match count, and the elapsed time. Turning it on costs a few percent of throughput and is almost always worth it, because most performance investigations start with a slow query nobody knew was being issued.'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'Getting started with the CLI'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
   &lt;span class="s1"&gt;'The command line client connects over the MySQL wire protocol, so any MySQL client works and you do not need to install anything special. Point it at port 9306 and you get an interactive shell. The shell understands the usual conveniences: history, tab completion of table names, and vertical output when a row is too wide for the terminal. Start by listing tables, then look at one with SHOW CREATE TABLE. The output is the exact statement that would recreate the table, including every option that was applied implicitly, which makes it the fastest way to find out what a table actually does rather than what someone documented two years ago. Bulk loading from the shell is possible but rarely what you want. For anything above a few thousand rows, use the HTTP bulk endpoint or one of the log shipper integrations, both of which batch and retry for you.'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'TLS and certificates for the HTTP API'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
   &lt;span class="s1"&gt;'The HTTP API can be served over TLS. You supply a certificate, a private key, and optionally a chain file, and the listener starts speaking HTTPS instead of HTTP. Clients that present a certificate of their own can be authenticated by it, which is the usual way to lock an internal API down without putting a password in every config file. Certificates for the HTTP API come from wherever your organisation gets certificates: a public authority, an internal authority, or an automated issuer. The file format is PEM. Both the certificate and the key must be readable by the user the server runs as, and the key must not be world readable or the listener refuses to start. Debugging TLS problems is mostly about reading the handshake. A client that reports an unknown authority is missing the chain. A client that reports a hostname mismatch is connecting by an address that is not in the certificate. A client that hangs is usually talking TLS to a plaintext port.'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now ask a question whose answer lives in the runbook's last section, once per strategy:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;knn_dist&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;docs&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;knn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;v_truncate&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'how do I rotate the TLS certificate used for replication'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;knn_dist&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;docs&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;knn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;v_mean&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'how do I rotate the TLS certificate used for replication'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;knn_dist&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;docs&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;knn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;v_sentence&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'how do I rotate the TLS certificate used for replication'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Strategy&lt;/th&gt;
&lt;th&gt;1st result&lt;/th&gt;
&lt;th&gt;2nd result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;truncate&lt;/code&gt; (default)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;TLS and certificates for the HTTP API&lt;/strong&gt; — 0.762&lt;/td&gt;
&lt;td&gt;Backup and restore runbook — 0.936&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;mean&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Backup and restore runbook — 0.656&lt;/td&gt;
&lt;td&gt;TLS and certificates for the HTTP API — 0.762&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;sentence&lt;/code&gt;, 128 tokens, 32 overlap&lt;/td&gt;
&lt;td&gt;Backup and restore runbook — &lt;strong&gt;0.254&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;TLS and certificates for the HTTP API — 0.700&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;With &lt;code&gt;truncate&lt;/code&gt;, the document that actually answers the question loses to a decoy that merely &lt;em&gt;looks&lt;/em&gt; like it is about certificates. The runbook's single vector was built from its opening pages on backup schedules and restore drills, because that is all the model was allowed to read.&lt;/p&gt;

&lt;p&gt;With &lt;code&gt;sentence&lt;/code&gt; chunking, the runbook is stored as nine vectors instead of one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;LENGTH&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;v_sentence&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;chunks&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;docs&lt;/span&gt; &lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="k"&gt;ASC&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;+------+---------------------------------------+--------+
| id   | title                                 | chunks |
+------+---------------------------------------+--------+
|    1 | Backup and restore runbook            |      9 |
|    2 | Monitoring and alerting guide         |      2 |
|    3 | Getting started with the CLI          |      2 |
|    4 | TLS and certificates for the HTTP API |      2 |
+------+---------------------------------------+--------+
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One of those nine is the certificate-rotation paragraph. It matches the query almost exactly, so the document wins by a wide margin: 0.254 against 0.700.&lt;/p&gt;

&lt;h2&gt;
  
  
  More about the chunking strategies
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Strategy&lt;/th&gt;
&lt;th&gt;Vectors per document&lt;/th&gt;
&lt;th&gt;Column type&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;truncate&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;&lt;code&gt;float_vector&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Embeds as much as fits the model's window, drops the rest. The only mode available in older versions, and still the default.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;mean&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;&lt;code&gt;float_vector&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Splits the whole document, embeds every piece, averages them into one vector.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;fixed&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;N&lt;/td&gt;
&lt;td&gt;&lt;code&gt;float_vector_array&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Fixed windows of &lt;code&gt;max_tokens&lt;/code&gt; tokens.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;recursive&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;N&lt;/td&gt;
&lt;td&gt;&lt;code&gt;float_vector_array&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Splits on a separator hierarchy — paragraph, then line, then sentence, then space — keeping each piece within &lt;code&gt;max_tokens&lt;/code&gt;.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;sentence&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;N&lt;/td&gt;
&lt;td&gt;&lt;code&gt;float_vector_array&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Sentence boundaries (&lt;a href="https://mnt.cr/go/gTMHqS" rel="noopener noreferrer"&gt;Unicode UAX #29&lt;/a&gt;), packed up to &lt;code&gt;max_tokens&lt;/code&gt;.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The important distinction is not how the text is cut. It is what a match &lt;em&gt;means&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;With one vector per document, search asks: &lt;strong&gt;"is this document, as a whole, similar to the query?"&lt;/strong&gt; A single relevant paragraph is diluted by everything around it, and a document that covers five topics ends up not really matching any of them.&lt;/p&gt;

&lt;p&gt;With one vector per chunk, search asks: &lt;strong&gt;"does this document &lt;em&gt;contain&lt;/em&gt; something similar?"&lt;/strong&gt; Each chunk competes on its own merits, and Manticore returns the document once, scored by its best chunk.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;truncate&lt;/code&gt; — keep it when your documents are short
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;title&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="n"&gt;float_vector&lt;/span&gt; &lt;span class="n"&gt;knn_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'hnsw'&lt;/span&gt; &lt;span class="n"&gt;hnsw_similarity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'cosine'&lt;/span&gt;
  &lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'Xenova/all-MiniLM-L6-v2'&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'title'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is what you already have; &lt;code&gt;chunk_strategy='truncate'&lt;/code&gt; is the default and you never have to write it. It's the right choice — and the fastest, and the smallest — whenever your text genuinely fits the model's window: product titles, short descriptions, tags, chat messages, log lines, search queries, commit subjects.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How much fits?&lt;/strong&gt; More than most people assume, and less than they hope. &lt;code&gt;all-MiniLM-L6-v2&lt;/code&gt; takes 512 tokens, roughly 380 English words. &lt;code&gt;text-embedding-3-small&lt;/code&gt; takes 8,192. If your 95th-percentile document is comfortably under the limit, stop reading and keep &lt;code&gt;truncate&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When it hurts:&lt;/strong&gt; anything long-form. Documentation pages, knowledge-base articles, contracts, transcripts, email threads, wiki pages, README files, incident postmortems.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;mean&lt;/code&gt; — one vector, but the whole document
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="n"&gt;float_vector&lt;/span&gt; &lt;span class="n"&gt;knn_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'hnsw'&lt;/span&gt; &lt;span class="n"&gt;hnsw_similarity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'cosine'&lt;/span&gt;
  &lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'Xenova/all-MiniLM-L6-v2'&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'title,content'&lt;/span&gt;
  &lt;span class="n"&gt;chunk_strategy&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'mean'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Manticore splits the document, embeds every chunk, and averages the chunk vectors into a single normalized vector. Storage and search cost are identical to &lt;code&gt;truncate&lt;/code&gt; — one vector per document, one HNSW node — but nothing is thrown away.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use it when:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You want the tail to count but can't afford more vectors — a very large corpus where index RAM is the binding constraint.&lt;/li&gt;
&lt;li&gt;The column is a plain &lt;code&gt;float_vector&lt;/code&gt; and you can't change the type (for example you're adding the column to an existing table with &lt;code&gt;ALTER&lt;/code&gt;, which multi-vector strategies don't support).&lt;/li&gt;
&lt;li&gt;Your documents are &lt;em&gt;about one thing&lt;/em&gt;, just long. A single product's full description, one recipe, one job posting.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Do not use it when&lt;/strong&gt; a document covers several unrelated topics. Averaging a legal contract's indemnity clause with its payment terms produces a vector that sits between them and is close to neither. In our benchmark below, &lt;code&gt;mean&lt;/code&gt; recovered about a third of the gap that chunking closes — a real improvement, and clearly not the same thing.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;fixed&lt;/code&gt; — predictable, cheapest to reason about
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="n"&gt;float_vector_array&lt;/span&gt; &lt;span class="n"&gt;knn_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'hnsw'&lt;/span&gt; &lt;span class="n"&gt;hnsw_similarity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'cosine'&lt;/span&gt;
  &lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'Xenova/all-MiniLM-L6-v2'&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'content'&lt;/span&gt;
  &lt;span class="n"&gt;chunk_strategy&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'fixed'&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'256'&lt;/span&gt; &lt;span class="n"&gt;overlap_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'32'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Cut every &lt;code&gt;max_tokens&lt;/code&gt; tokens, no matter what the text is doing at that point. Chunk count is a straight function of document length, so index size is predictable before you load anything.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use it when&lt;/strong&gt; the text has no reliable structure to exploit: OCR output, scraped HTML that lost its paragraphs, machine transcripts without punctuation, log dumps, minified content. Also a fine default when you simply want the cheapest thing that stops truncation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The cost:&lt;/strong&gt; a boundary can land mid-sentence, and a chunk that begins in the middle of a thought embeds badly. That is exactly what &lt;code&gt;overlap_tokens&lt;/code&gt; is for — see below.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;recursive&lt;/code&gt; — the best general default for prose
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="n"&gt;float_vector_array&lt;/span&gt; &lt;span class="n"&gt;knn_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'hnsw'&lt;/span&gt; &lt;span class="n"&gt;hnsw_similarity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'cosine'&lt;/span&gt;
  &lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'Xenova/all-MiniLM-L6-v2'&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'title,content'&lt;/span&gt;
  &lt;span class="n"&gt;chunk_strategy&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'recursive'&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'256'&lt;/span&gt; &lt;span class="n"&gt;overlap_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'32'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same token budget as &lt;code&gt;fixed&lt;/code&gt;, but each cut is pulled back to the nearest natural boundary: a blank line first, then a line break, then a sentence end, then a space. A chunk stops where the text stops, not where the counter runs out. The boundary is never dragged back past the midpoint of the chunk, so you don't get a stream of tiny fragments.&lt;/p&gt;

&lt;p&gt;If you have used LangChain's &lt;code&gt;RecursiveCharacterTextSplitter&lt;/code&gt;, this is the same idea, except it runs inside the database on the model's real tokens instead of characters, and there is nothing to install.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use it for:&lt;/strong&gt; Markdown and HTML documentation, wiki pages, knowledge bases, blog posts, README files, structured reports — anything written by a human in paragraphs. This scored highest on deep content in our benchmark.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;sentence&lt;/code&gt; — when a chunk must be a complete thought
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="n"&gt;float_vector_array&lt;/span&gt; &lt;span class="n"&gt;knn_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'hnsw'&lt;/span&gt; &lt;span class="n"&gt;hnsw_similarity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'cosine'&lt;/span&gt;
  &lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'Xenova/all-MiniLM-L6-v2'&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'content'&lt;/span&gt;
  &lt;span class="n"&gt;chunk_strategy&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'sentence'&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'256'&lt;/span&gt; &lt;span class="n"&gt;overlap_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'32'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Detects sentence boundaries with &lt;a href="https://mnt.cr/go/gTMHqS" rel="noopener noreferrer"&gt;Unicode UAX #29&lt;/a&gt;, then greedily packs whole sentences until the token budget is reached. A chunk never starts or ends mid-sentence. A single sentence longer than the budget is split by the token window, as a last resort.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use it for:&lt;/strong&gt; support tickets and email threads, chat and meeting transcripts, legal and policy text, news, customer reviews, medical and scientific abstracts — anything where a fragment of a sentence changes or destroys the meaning. It is also the strategy to pick when chunks will be fed to an LLM afterwards, because a chunk that ends mid-clause reads badly in a prompt.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;sentence&lt;/code&gt; is a little more conservative than &lt;code&gt;recursive&lt;/code&gt;: it produced fewer, cleaner chunks in our tests and scored about the same on recall@5.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three knobs
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight conf"&gt;&lt;code&gt;&lt;span class="n"&gt;chunk_strategy&lt;/span&gt;  = &lt;span class="n"&gt;truncate&lt;/span&gt; | &lt;span class="n"&gt;mean&lt;/span&gt; | &lt;span class="n"&gt;fixed&lt;/span&gt; | &lt;span class="n"&gt;recursive&lt;/span&gt; | &lt;span class="n"&gt;sentence&lt;/span&gt;
&lt;span class="n"&gt;max_tokens&lt;/span&gt;      = &lt;span class="n"&gt;chunk&lt;/span&gt; &lt;span class="n"&gt;size&lt;/span&gt; &lt;span class="n"&gt;in&lt;/span&gt; &lt;span class="n"&gt;tokens&lt;/span&gt;; &lt;span class="m"&gt;0&lt;/span&gt; (&lt;span class="n"&gt;default&lt;/span&gt;) = &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="err"&gt;'&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="n"&gt;own&lt;/span&gt; &lt;span class="n"&gt;limit&lt;/span&gt;
&lt;span class="n"&gt;overlap_tokens&lt;/span&gt;  = &lt;span class="n"&gt;tokens&lt;/span&gt; &lt;span class="n"&gt;shared&lt;/span&gt; &lt;span class="n"&gt;between&lt;/span&gt; &lt;span class="n"&gt;consecutive&lt;/span&gt; &lt;span class="n"&gt;chunks&lt;/span&gt;; &lt;span class="n"&gt;needs&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="n"&gt;non&lt;/span&gt;-&lt;span class="n"&gt;zero&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;
&lt;span class="n"&gt;max_chunks&lt;/span&gt;      = &lt;span class="n"&gt;ceiling&lt;/span&gt; &lt;span class="n"&gt;on&lt;/span&gt; &lt;span class="n"&gt;vectors&lt;/span&gt; &lt;span class="n"&gt;per&lt;/span&gt; &lt;span class="n"&gt;document&lt;/span&gt;; &lt;span class="m"&gt;0&lt;/span&gt; (&lt;span class="n"&gt;default&lt;/span&gt;) = &lt;span class="n"&gt;unlimited&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;&lt;code&gt;max_tokens&lt;/code&gt;&lt;/strong&gt; is capped at what the model can actually accept — ask for 4,096 on a 512-token model and you still get 512, not an error. Smaller chunks mean sharper matches and more vectors; larger chunks mean more context per vector and fewer of them. For English prose, 128–512 covers almost every use case; we used 256 throughout the benchmark.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;overlap_tokens&lt;/code&gt;&lt;/strong&gt; repeats the tail of each chunk at the head of the next, so a sentence that straddles a boundary still appears intact somewhere. 10–20% of &lt;code&gt;max_tokens&lt;/code&gt; is the usual setting. Manticore guarantees forward progress: &lt;code&gt;fixed&lt;/code&gt; and &lt;code&gt;recursive&lt;/code&gt; cap the overlap at half the chunk size, and &lt;code&gt;sentence&lt;/code&gt; re-seeds the next chunk with at most &lt;code&gt;overlap_tokens&lt;/code&gt; worth of trailing whole sentences while always advancing by at least one sentence. It requires an explicit non-zero &lt;code&gt;max_tokens&lt;/code&gt; — overlap against "whatever the model's limit happens to be" isn't a meaningful setting, so Manticore rejects it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;max_chunks&lt;/code&gt;&lt;/strong&gt; limits the impact of unusually large documents. Without it, a 400-page PDF pasted into one row becomes thousands of HNSW nodes. With it, Manticore merges the overflow into the last kept chunk, then truncates it to the model's window when embedding:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- a ~600-token document, chunked at 64 tokens&lt;/span&gt;
&lt;span class="n"&gt;chunk_strategy&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'fixed'&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'64'&lt;/span&gt;                  &lt;span class="c1"&gt;-- 22 vectors&lt;/span&gt;
&lt;span class="n"&gt;chunk_strategy&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'fixed'&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'64'&lt;/span&gt; &lt;span class="n"&gt;max_chunks&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'3'&lt;/span&gt;   &lt;span class="c1"&gt;--  3 vectors&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use it as a guard rail against outliers, not as a way to save memory across the board.&lt;/p&gt;

&lt;h2&gt;
  
  
  What search looks like
&lt;/h2&gt;

&lt;p&gt;Nothing about your query changes from before. There is no chunk table, no nested field, no join, no &lt;code&gt;GROUP BY&lt;/code&gt;. Here is the complete example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;DROP&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;IF&lt;/span&gt; &lt;span class="k"&gt;EXISTS&lt;/span&gt; &lt;span class="n"&gt;notes&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;notes&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;title&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;body&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;chunks&lt;/span&gt; &lt;span class="n"&gt;float_vector_array&lt;/span&gt; &lt;span class="n"&gt;knn_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'hnsw'&lt;/span&gt; &lt;span class="n"&gt;hnsw_similarity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'cosine'&lt;/span&gt;
    &lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'Xenova/all-MiniLM-L6-v2'&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'title,body'&lt;/span&gt;
    &lt;span class="n"&gt;chunk_strategy&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'sentence'&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'32'&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;INSERT&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;notes&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;VALUES&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'Certificate rotation'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
   &lt;span class="s1"&gt;'The replication certificate is not the one the HTTP API uses. Generate the new key on the node being rotated and sign it with the cluster authority. Update the paths and reload, one node at a time, confirming every peer reports as synced before you move on.'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'Disk pressure'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
   &lt;span class="s1"&gt;'When a data directory crosses eighty percent the merge scheduler stops compacting and the node starts refusing writes. Free space first, then trigger a manual OPTIMIZE. Adding a disk without draining the queue only postpones the problem.'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'Slow queries'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
   &lt;span class="s1"&gt;'Turn the query log on before guessing. Most investigations end at a single query nobody knew was being issued, usually one that sorts on an unindexed attribute over the whole table.'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;knn_dist&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;notes&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;knn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'how do I replace an expiring certificate on every node'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;+------+----------------------+------------+
| id   | title                | knn_dist() |
+------+----------------------+------------+
|    1 | Certificate rotation | 0.51039070 |
|    2 | Disk pressure        | 0.91703475 |
|    3 | Slow queries         | 1.02606630 |
+------+----------------------+------------+
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;max_tokens='32'&lt;/code&gt; is small on purpose here, so that these short notes actually split and you can see the multi-vector behaviour on a toy dataset. &lt;code&gt;LENGTH()&lt;/code&gt; on the vector column shows how each document was divided:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;LENGTH&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;notes&lt;/span&gt; &lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt; &lt;span class="k"&gt;LIMIT&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;+------+----------------------+------+
| id   | title                | n    |
+------+----------------------+------+
|    1 | Certificate rotation |    2 |
|    2 | Disk pressure        |    2 |
|    3 | Slow queries         |    2 |
+------+----------------------+------+
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Six vectors, three rows back. Search follows the rules described in &lt;a href="https://mnt.cr/go/7D6PF2" rel="noopener noreferrer"&gt;Multiple vectors per document&lt;/a&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A document matches if &lt;strong&gt;any&lt;/strong&gt; of its vectors is near the query vector.&lt;/li&gt;
&lt;li&gt;Manticore returns each match &lt;strong&gt;exactly once&lt;/strong&gt;. &lt;code&gt;knn_dist()&lt;/code&gt; is the distance to its &lt;strong&gt;closest&lt;/strong&gt; chunk.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;k&lt;/code&gt; counts documents&lt;/strong&gt;, not vectors. &lt;code&gt;knn(chunks, 3, ...)&lt;/code&gt; means three documents.&lt;/li&gt;
&lt;li&gt;A document with no vectors is never returned.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The same query over HTTP:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;POST&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;/search&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"table"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"notes"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"knn"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"field"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"chunks"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"query"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"how do I replace an expiring certificate on every node"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"k"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"_source"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"title"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;...&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"_score"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"_knn_dist"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.51039070&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"_source"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"title"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Certificate rotation"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;...&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Everything else on the KNN page keeps working as before: filtering, prefilter and postfilter strategies, &lt;a href="https://mnt.cr/go/p4zmAK" rel="noopener noreferrer"&gt;quantization&lt;/a&gt;, &lt;a href="https://mnt.cr/go/W4eDvC" rel="noopener noreferrer"&gt;early termination&lt;/a&gt;, and rescoring.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does it actually help? Numbers on our own manual
&lt;/h2&gt;

&lt;p&gt;We tested the feature on the &lt;strong&gt;Manticore English manual&lt;/strong&gt; — 189 pages and about 298,000 words, ranging from a two-paragraph note to a 39,000-word changelog.&lt;/p&gt;

&lt;p&gt;The query set is generated mechanically, not hand-picked. For every page we took its section headings, kept only headings that are unique across the whole manual, and split them in two:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Deep-content queries (419)&lt;/strong&gt; — headings that appear &lt;em&gt;after&lt;/em&gt; the first ~1,200 characters of their page. That is a deliberately conservative line: the model's window is 512 tokens, roughly 2,000 characters, so a few of these still point at text &lt;code&gt;truncate&lt;/code&gt; can partly see. The gap below is therefore an understatement, not an exaggeration.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Head-content queries (88)&lt;/strong&gt; — headings inside the first ~1,200 characters. The control group: content &lt;code&gt;truncate&lt;/code&gt; can already see.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A query is a hit if KNN returns the page the heading came from, within the top &lt;em&gt;k&lt;/em&gt;. Model: &lt;code&gt;Xenova/all-MiniLM-L6-v2&lt;/code&gt; (384 dims, 512-token window) running on Manticore's &lt;a href="https://mnt.cr/go/JWrvN3" rel="noopener noreferrer"&gt;ONNX backend&lt;/a&gt;. Hardware: 32 threads. &lt;code&gt;max_tokens='256'&lt;/code&gt;, &lt;code&gt;overlap_tokens='32'&lt;/code&gt; for the multi-vector strategies. Quality numbers are deterministic for a given index; the timings are a single run per strategy on an otherwise idle box.&lt;/p&gt;

&lt;h3&gt;
  
  
  Deep content — what chunking is for
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Strategy&lt;/th&gt;
&lt;th&gt;Vectors&lt;/th&gt;
&lt;th&gt;Ingest&lt;/th&gt;
&lt;th&gt;Index RAM&lt;/th&gt;
&lt;th&gt;hit@1&lt;/th&gt;
&lt;th&gt;hit@5&lt;/th&gt;
&lt;th&gt;hit@10&lt;/th&gt;
&lt;th&gt;MRR&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;truncate&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;189&lt;/td&gt;
&lt;td&gt;21 s&lt;/td&gt;
&lt;td&gt;4.2 MB&lt;/td&gt;
&lt;td&gt;33.7%&lt;/td&gt;
&lt;td&gt;55.1%&lt;/td&gt;
&lt;td&gt;63.2%&lt;/td&gt;
&lt;td&gt;0.44&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;mean&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;189&lt;/td&gt;
&lt;td&gt;72 s&lt;/td&gt;
&lt;td&gt;4.2 MB&lt;/td&gt;
&lt;td&gt;43.9%&lt;/td&gt;
&lt;td&gt;65.2%&lt;/td&gt;
&lt;td&gt;74.7%&lt;/td&gt;
&lt;td&gt;0.54&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;fixed&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;3,430&lt;/td&gt;
&lt;td&gt;73 s&lt;/td&gt;
&lt;td&gt;9.5 MB&lt;/td&gt;
&lt;td&gt;56.3%&lt;/td&gt;
&lt;td&gt;81.1%&lt;/td&gt;
&lt;td&gt;86.2%&lt;/td&gt;
&lt;td&gt;0.68&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;recursive&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;4,664&lt;/td&gt;
&lt;td&gt;86 s&lt;/td&gt;
&lt;td&gt;11.7 MB&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;58.7%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;83.3%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;89.5%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.70&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;sentence&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;4,041&lt;/td&gt;
&lt;td&gt;79 s&lt;/td&gt;
&lt;td&gt;10.6 MB&lt;/td&gt;
&lt;td&gt;55.4%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;83.5%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;89.0%&lt;/td&gt;
&lt;td&gt;0.68&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Chunking turns a coin flip into a working search. &lt;strong&gt;recall@5 goes from 55.1% to 83.3%&lt;/strong&gt;, and the rank of the right answer improves just as much — MRR 0.44 → 0.70. Of the queries &lt;code&gt;truncate&lt;/code&gt; could not answer in the top 5 at all, &lt;code&gt;recursive&lt;/code&gt; recovers roughly two thirds.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;mean&lt;/code&gt; lands where you would expect: it recovers about a third of the gap for free, because it costs exactly nothing extra to store or search.&lt;/p&gt;

&lt;h3&gt;
  
  
  Head content — the control group
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Strategy&lt;/th&gt;
&lt;th&gt;hit@1&lt;/th&gt;
&lt;th&gt;hit@5&lt;/th&gt;
&lt;th&gt;MRR&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;truncate&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;65.9%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;86.4%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.74&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;mean&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;59.1%&lt;/td&gt;
&lt;td&gt;83.0%&lt;/td&gt;
&lt;td&gt;0.69&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;fixed&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;60.2%&lt;/td&gt;
&lt;td&gt;83.0%&lt;/td&gt;
&lt;td&gt;0.71&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;recursive&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;58.0%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;86.4%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0.70&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;sentence&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;56.8%&lt;/td&gt;
&lt;td&gt;85.2%&lt;/td&gt;
&lt;td&gt;0.69&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For completeness, the control group is worth reading carefully. For content that the model could already see, &lt;strong&gt;&lt;code&gt;truncate&lt;/code&gt; is still the most precise at rank 1&lt;/strong&gt; — 65.9% against 58.0% for &lt;code&gt;recursive&lt;/code&gt;. A whole-document vector carries the page's overall topic, and when the query is about the page's opening subject, that context helps.&lt;/p&gt;

&lt;p&gt;By rank 5 the difference is gone: &lt;code&gt;recursive&lt;/code&gt; matches &lt;code&gt;truncate&lt;/code&gt; exactly at 86.4%. So the trade is a few points of top-1 precision on content near the beginning, in exchange for +28 points of recall on everything else. For a documentation search, a help center, or any RAG retriever that feeds 5–10 passages to an LLM, that is not a close call.&lt;/p&gt;

&lt;h3&gt;
  
  
  Cost
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Index RAM&lt;/strong&gt;: 4.2 MB → 11.7 MB, about 2.5×, for ~25× as many vectors. Vectors are only part of what an RT table stores. The HNSW graph over those vectors also takes longer to build during chunk saves and &lt;code&gt;OPTIMIZE&lt;/code&gt;, though Manticore &lt;a href="https://mnt.cr/go/If4DTE" rel="noopener noreferrer"&gt;builds it across all your cores&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Data load&lt;/strong&gt;: 21 s → 86 s for 189 documents. Chunking means embedding the whole corpus instead of the first 380 words of each document, and the time scales with it. This is embedding cost, not chunking cost — the splitting itself is not measurable next to inference.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Query response time&lt;/strong&gt;: 6.3 ms → 8.5 ms at p50. HNSW handles 4,664 vectors about as easily as 189 — see &lt;a href="https://mnt.cr/go/IAAbBb" rel="noopener noreferrer"&gt;2-pass HNSW, batched distances and AVX-512&lt;/a&gt; for what carries that.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you use a paid embedding API, read that ingest number as a bill: chunking sends your whole corpus to the model instead of the head of each document, and you pay for every token of it. Local ONNX models have no per-token cost, which is a large part of why we made them &lt;a href="https://mnt.cr/go/JWrvN3" rel="noopener noreferrer"&gt;fast&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recommendations for choosing a chunking strategy
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Your data&lt;/th&gt;
&lt;th&gt;Start with&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Titles, names, short descriptions, tags, log lines&lt;/td&gt;
&lt;td&gt;&lt;code&gt;truncate&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Long but single-topic; or RAM is the hard limit; or the column is an existing &lt;code&gt;float_vector&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;&lt;code&gt;mean&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Documentation, wikis, knowledge bases, articles, READMEs&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;recursive&lt;/code&gt;, &lt;code&gt;max_tokens&lt;/code&gt; 128–256&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Support tickets, email, transcripts, legal text, reviews&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;sentence&lt;/code&gt;, &lt;code&gt;max_tokens&lt;/code&gt; 128–256&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OCR, scraped HTML, machine transcripts, unstructured dumps&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;fixed&lt;/code&gt;, &lt;code&gt;max_tokens&lt;/code&gt; 256, plus overlap&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chunks will be passed to an LLM as context&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;sentence&lt;/code&gt;, &lt;code&gt;max_tokens&lt;/code&gt; 384–512&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Overlap is deliberately absent from most of those: our sweep below could not measure a benefit from it on structured prose, and it costs vectors. Add it when a thought routinely straddles a boundary — unstructured transcripts, OCR, long narrative without paragraph breaks.&lt;/p&gt;

&lt;h3&gt;
  
  
  How big should a chunk be?
&lt;/h3&gt;

&lt;p&gt;Chunk size is the setting that actually affects your results. The trade is direct: &lt;strong&gt;a smaller chunk is a sharper match on one idea, a larger chunk carries more context but dilutes each idea inside it.&lt;/strong&gt; A paragraph buried in a long document only becomes findable once the chunk size is small enough to give it a vector of its own.&lt;/p&gt;

&lt;p&gt;We ran another test: &lt;code&gt;recursive&lt;/code&gt; on the same 189-page manual, three chunk sizes × three overlap settings, and the same 419 deep queries. Two trends stand out: quality rises as chunks get smaller, while overlap adds cost without improving quality much.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmnt.cr%2Fgo%2FluK0mM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmnt.cr%2Fgo%2FluK0mM" alt="Retrieval quality and index RAM against chunk size, for three overlap settings"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Note that the Y axis starts at 78%, not zero — the whole spread is about six points, so a zero-based axis would flatten it into a straight line. The numbers behind the chart:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;code&gt;max_tokens&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;&lt;code&gt;overlap_tokens&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;Vectors&lt;/th&gt;
&lt;th&gt;Index RAM&lt;/th&gt;
&lt;th&gt;deep hit@5&lt;/th&gt;
&lt;th&gt;deep MRR&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;128&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;8,256&lt;/td&gt;
&lt;td&gt;17.3 MB&lt;/td&gt;
&lt;td&gt;85.2%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.718&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;128&lt;/td&gt;
&lt;td&gt;13&lt;/td&gt;
&lt;td&gt;9,328&lt;/td&gt;
&lt;td&gt;18.7 MB&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;85.7%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0.705&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;128&lt;/td&gt;
&lt;td&gt;32&lt;/td&gt;
&lt;td&gt;11,623&lt;/td&gt;
&lt;td&gt;22.6 MB&lt;/td&gt;
&lt;td&gt;85.2%&lt;/td&gt;
&lt;td&gt;0.694&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;256&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;3,984&lt;/td&gt;
&lt;td&gt;10.0 MB&lt;/td&gt;
&lt;td&gt;83.1%&lt;/td&gt;
&lt;td&gt;0.657&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;256&lt;/td&gt;
&lt;td&gt;26&lt;/td&gt;
&lt;td&gt;4,525&lt;/td&gt;
&lt;td&gt;10.9 MB&lt;/td&gt;
&lt;td&gt;84.5%&lt;/td&gt;
&lt;td&gt;0.681&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;256&lt;/td&gt;
&lt;td&gt;64&lt;/td&gt;
&lt;td&gt;5,569&lt;/td&gt;
&lt;td&gt;12.5 MB&lt;/td&gt;
&lt;td&gt;83.5%&lt;/td&gt;
&lt;td&gt;0.689&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;512&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;1,973&lt;/td&gt;
&lt;td&gt;6.9 MB&lt;/td&gt;
&lt;td&gt;80.2%&lt;/td&gt;
&lt;td&gt;0.655&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;512&lt;/td&gt;
&lt;td&gt;51&lt;/td&gt;
&lt;td&gt;2,191&lt;/td&gt;
&lt;td&gt;7.2 MB&lt;/td&gt;
&lt;td&gt;79.2%&lt;/td&gt;
&lt;td&gt;0.651&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;512&lt;/td&gt;
&lt;td&gt;128&lt;/td&gt;
&lt;td&gt;2,666&lt;/td&gt;
&lt;td&gt;8.0 MB&lt;/td&gt;
&lt;td&gt;79.7%&lt;/td&gt;
&lt;td&gt;0.660&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;What we see:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Smaller chunks win, consistently.&lt;/strong&gt; Going from 512 to 128 tokens buys about five points of recall@5 (80.2% → 85.2%) and a large jump in ranking quality (MRR 0.655 → 0.718). It costs 4× the vectors and 2.5× the index RAM. Below 128 the chunks stop containing a whole thought, so this is not a slope you ride forever — but on long technical prose, 128–256 beat 512 every time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Overlap did essentially nothing for quality, and was not free.&lt;/strong&gt; At 128 tokens, going from no overlap to 25% overlap moved recall@5 from 85.2% to 85.2% while adding 41% more vectors and 5 MB of RAM. The pattern holds at every size: the spread across overlap settings (±1.5 points) is within the noise of a 419-query set, while the cost is not. This lines up with Chroma's &lt;a href="https://mnt.cr/go/J74ij5" rel="noopener noreferrer"&gt;chunking evaluation&lt;/a&gt;, where plain recursive splitting at 200 tokens &lt;strong&gt;with no overlap&lt;/strong&gt; scored 88.1% recall — within a few points of an LLM-driven splitter at 91.9% — and it is the opposite of the "always use 10–20% overlap" advice you will read in most RAG guides.&lt;/p&gt;

&lt;p&gt;The honest caveat: this is one corpus, one model, and queries that look like section headings. Overlap earns its keep when a single fact routinely straddles a boundary — long unbroken narrative, transcripts without structure — and &lt;code&gt;recursive&lt;/code&gt; already snaps cuts to paragraph and sentence boundaries, which does much of the same work. So treat "start at 128–256 with no overlap, add overlap only if you can measure it helping" as the default, and check it on your own data with the recipe below.&lt;/p&gt;

&lt;h3&gt;
  
  
  Comparing settings
&lt;/h3&gt;

&lt;p&gt;You do not have to guess, and you do not need two tables. A table can carry &lt;strong&gt;several model-backed vector columns&lt;/strong&gt;, each with its own strategy, all filled from the same fields on the same &lt;code&gt;INSERT&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;ab&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;title&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;body&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;sent_256&lt;/span&gt; &lt;span class="n"&gt;float_vector_array&lt;/span&gt; &lt;span class="n"&gt;knn_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'hnsw'&lt;/span&gt; &lt;span class="n"&gt;hnsw_similarity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'cosine'&lt;/span&gt;
    &lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'Xenova/all-MiniLM-L6-v2'&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'title,body'&lt;/span&gt;
    &lt;span class="n"&gt;chunk_strategy&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'sentence'&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'256'&lt;/span&gt; &lt;span class="n"&gt;overlap_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'32'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;rec_128&lt;/span&gt; &lt;span class="n"&gt;float_vector_array&lt;/span&gt; &lt;span class="n"&gt;knn_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'hnsw'&lt;/span&gt; &lt;span class="n"&gt;hnsw_similarity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'cosine'&lt;/span&gt;
    &lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'Xenova/all-MiniLM-L6-v2'&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'title,body'&lt;/span&gt;
    &lt;span class="n"&gt;chunk_strategy&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'recursive'&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'128'&lt;/span&gt; &lt;span class="n"&gt;overlap_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'16'&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Load your corpus once, then run the same query against each column and compare. For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;LENGTH&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sent_256&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;sent_chunks&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;LENGTH&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rec_128&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;rec_chunks&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;ab&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;knn_dist&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;ab&lt;/span&gt; &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;knn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sent_256&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'how do I rotate the replication certificate'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;knn_dist&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;ab&lt;/span&gt; &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;knn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rec_128&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'how do I rotate the replication certificate'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On a short runbook whose certificate section sits at the end, &lt;code&gt;sentence&lt;/code&gt;/256 fits the whole document in a single chunk and answers at distance &lt;strong&gt;0.515&lt;/strong&gt;; &lt;code&gt;recursive&lt;/code&gt;/128 splits it in two, isolates the certificate paragraph, and answers at &lt;strong&gt;0.310&lt;/strong&gt;. Same row, same model, same query — only the chunk size differs.&lt;/p&gt;

&lt;p&gt;Build a dataset of real queries with answers you trust — even 50 is enough — and compare recall@5 across two or three columns, exactly as we did on the manual above. Then drop the losing column with &lt;code&gt;ALTER TABLE ... DROP COLUMN&lt;/code&gt; and keep the winner.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recipes
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Documentation and help center search.&lt;/strong&gt; Long Markdown pages, users asking questions in their own words. Chunk on structure and search across it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;docs&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;url&lt;/span&gt; &lt;span class="n"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;title&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;body&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;chunks&lt;/span&gt; &lt;span class="n"&gt;float_vector_array&lt;/span&gt; &lt;span class="n"&gt;knn_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'hnsw'&lt;/span&gt; &lt;span class="n"&gt;hnsw_similarity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'cosine'&lt;/span&gt;
    &lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'Xenova/all-MiniLM-L6-v2'&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'title,body'&lt;/span&gt;
    &lt;span class="n"&gt;chunk_strategy&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'recursive'&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'192'&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Note &lt;code&gt;from='title,body'&lt;/code&gt;: the fields are joined before chunking, so the page title lands in the first chunk and gives it context. For a worked end-to-end example of this shape, see &lt;a href="https://mnt.cr/go/euc1eN" rel="noopener noreferrer"&gt;Vector search on GitHub&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Support tickets and email threads.&lt;/strong&gt; A thread is a sequence of complete messages; cutting one mid-sentence loses the fact you need. Keep the chunk count bounded, because threads have no natural length limit:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;tickets&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;ticket_id&lt;/span&gt; &lt;span class="nb"&gt;bigint&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;customer&lt;/span&gt; &lt;span class="n"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="n"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;thread&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;chunks&lt;/span&gt; &lt;span class="n"&gt;float_vector_array&lt;/span&gt; &lt;span class="n"&gt;knn_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'hnsw'&lt;/span&gt; &lt;span class="n"&gt;hnsw_similarity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'cosine'&lt;/span&gt;
    &lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'Xenova/all-MiniLM-L6-v2'&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'thread'&lt;/span&gt;
    &lt;span class="n"&gt;chunk_strategy&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'sentence'&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'256'&lt;/span&gt; &lt;span class="n"&gt;overlap_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'32'&lt;/span&gt; &lt;span class="n"&gt;max_chunks&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'64'&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;ticket_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;knn_dist&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;tickets&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;knn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'customer was charged twice after upgrading'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'closed'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Filtering works exactly as it does for a single-vector column.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Contracts and policy documents.&lt;/strong&gt; Clause-level retrieval is the entire point — nobody wants "the contract" back, they want the indemnity clause. Smaller chunks, generous overlap:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;chunks&lt;/span&gt; &lt;span class="n"&gt;float_vector_array&lt;/span&gt; &lt;span class="n"&gt;knn_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'hnsw'&lt;/span&gt; &lt;span class="n"&gt;hnsw_similarity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'cosine'&lt;/span&gt;
  &lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'Xenova/all-MiniLM-L6-v2'&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'body'&lt;/span&gt;
  &lt;span class="n"&gt;chunk_strategy&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'sentence'&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'128'&lt;/span&gt; &lt;span class="n"&gt;overlap_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'32'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Product catalog with long descriptions.&lt;/strong&gt; One product is one topic, and catalogs are large, so pay nothing extra:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;embedding&lt;/span&gt; &lt;span class="n"&gt;float_vector&lt;/span&gt; &lt;span class="n"&gt;knn_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'hnsw'&lt;/span&gt; &lt;span class="n"&gt;hnsw_similarity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'cosine'&lt;/span&gt;
  &lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'Xenova/all-MiniLM-L6-v2'&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'name,description'&lt;/span&gt;
  &lt;span class="n"&gt;chunk_strategy&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'mean'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;RAG: retrieval for an LLM.&lt;/strong&gt; Whatever you retrieve gets pasted into a prompt, so chunks should read as prose — this is the retrieval half of &lt;a href="https://mnt.cr/go/zcanB7" rel="noopener noreferrer"&gt;conversational search&lt;/a&gt;. Larger chunks, sentence boundaries, and ask for more of them:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;chunks&lt;/span&gt; &lt;span class="n"&gt;float_vector_array&lt;/span&gt; &lt;span class="n"&gt;knn_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'hnsw'&lt;/span&gt; &lt;span class="n"&gt;hnsw_similarity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'cosine'&lt;/span&gt;
  &lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'Xenova/all-MiniLM-L6-v2'&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'title,body'&lt;/span&gt;
  &lt;span class="n"&gt;chunk_strategy&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'sentence'&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'512'&lt;/span&gt; &lt;span class="n"&gt;overlap_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'64'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Adding chunking to a table you already have.&lt;/strong&gt; Multi-vector columns can't be added by &lt;code&gt;ALTER&lt;/code&gt; — existing rows have no vectors and there's no way to backfill them yet. A single-vector strategy can:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;docs&lt;/span&gt; &lt;span class="k"&gt;ADD&lt;/span&gt; &lt;span class="k"&gt;COLUMN&lt;/span&gt; &lt;span class="n"&gt;v2&lt;/span&gt; &lt;span class="n"&gt;float_vector&lt;/span&gt; &lt;span class="n"&gt;knn_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'hnsw'&lt;/span&gt; &lt;span class="n"&gt;hnsw_similarity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'cosine'&lt;/span&gt;
  &lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'Xenova/all-MiniLM-L6-v2'&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'title,body'&lt;/span&gt; &lt;span class="n"&gt;chunk_strategy&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'mean'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;docs&lt;/span&gt; &lt;span class="n"&gt;REBUILD&lt;/span&gt; &lt;span class="n"&gt;EMBEDDINGS&lt;/span&gt; &lt;span class="n"&gt;v2&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For a multi-vector column, create the new table with the column in place and reindex into it.&lt;/p&gt;

&lt;h2&gt;
  
  
  How other engines handle this
&lt;/h2&gt;

&lt;p&gt;Every vector engine now generates embeddings for you. Far fewer will &lt;em&gt;split&lt;/em&gt; your document before doing it — and of those, most make you assemble it out of pipeline stages.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Engine&lt;/th&gt;
&lt;th&gt;Embeds in-engine&lt;/th&gt;
&lt;th&gt;Chunks in-engine&lt;/th&gt;
&lt;th&gt;Strategies&lt;/th&gt;
&lt;th&gt;One row per document at search time&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Manticore Search&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes — local + OpenAI / Voyage / Jina&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;&lt;a href="https://mnt.cr/go/3aXp5K" rel="noopener noreferrer"&gt;Yes — &lt;code&gt;chunk_strategy&lt;/code&gt; on the vector column&lt;/a&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;truncate, mean, fixed, recursive, sentence&lt;/td&gt;
&lt;td&gt;Yes, native&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Elasticsearch&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes — inference endpoints&lt;/td&gt;
&lt;td&gt;&lt;a href="https://mnt.cr/go/MkrPN7" rel="noopener noreferrer"&gt;Yes&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;sentence (default), word, recursive (9.1+), none&lt;/td&gt;
&lt;td&gt;Yes — &lt;code&gt;semantic_text&lt;/code&gt; hides the chunks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;OpenSearch&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes — ML Commons&lt;/td&gt;
&lt;td&gt;&lt;a href="https://mnt.cr/go/7lTFni" rel="noopener noreferrer"&gt;Yes — separate ingest processor&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;fixed_token_length, fixed_char_length, delimiter&lt;/td&gt;
&lt;td&gt;Needs a nested field + nested query&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Vespa&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes — built-in embedders&lt;/td&gt;
&lt;td&gt;&lt;a href="https://mnt.cr/go/XwlZBv" rel="noopener noreferrer"&gt;Yes — indexing expression&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;fixed-length, sentence, custom&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Azure AI Search&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes — integrated vectorization&lt;/td&gt;
&lt;td&gt;&lt;a href="https://mnt.cr/go/iFze6f" rel="noopener noreferrer"&gt;Yes — Split skill in a skillset&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;pages (chars), sentences&lt;/td&gt;
&lt;td&gt;No — one row per chunk&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;PostgreSQL + pgai&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes — background worker&lt;/td&gt;
&lt;td&gt;&lt;a href="https://mnt.cr/go/KRJkJF" rel="noopener noreferrer"&gt;Yes&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;character, recursive character&lt;/td&gt;
&lt;td&gt;No — separate table, join and dedupe&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Milvus / Zilliz&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes — Function (2.6+)&lt;/td&gt;
&lt;td&gt;No — &lt;a href="https://mnt.cr/go/Cw8LEL" rel="noopener noreferrer"&gt;app-side&lt;/a&gt;
&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Qdrant&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes — Cloud Inference&lt;/td&gt;
&lt;td&gt;No — &lt;a href="https://mnt.cr/go/qSAAtd" rel="noopener noreferrer"&gt;app-side&lt;/a&gt;
&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Weaviate&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes — vectorizer modules&lt;/td&gt;
&lt;td&gt;No — &lt;a href="https://mnt.cr/go/dcwvFh" rel="noopener noreferrer"&gt;app-side&lt;/a&gt;
&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Meilisearch&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes — embedders&lt;/td&gt;
&lt;td&gt;No — &lt;a href="https://mnt.cr/go/jiVAam" rel="noopener noreferrer"&gt;app-side&lt;/a&gt;
&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Typesense&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No — &lt;a href="https://mnt.cr/go/rkcCr1" rel="noopener noreferrer"&gt;open request&lt;/a&gt;
&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Apache Solr&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes — &lt;a href="https://mnt.cr/go/p53RDj" rel="noopener noreferrer"&gt;LLM module&lt;/a&gt; (9.8+)&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Pinecone&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes — integrated inference&lt;/td&gt;
&lt;td&gt;No — &lt;a href="https://mnt.cr/go/DfWdRL" rel="noopener noreferrer"&gt;app-side&lt;/a&gt;
&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;MongoDB Atlas&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes — Automated Embedding&lt;/td&gt;
&lt;td&gt;No — &lt;a href="https://mnt.cr/go/ht2vzt" rel="noopener noreferrer"&gt;app-side&lt;/a&gt;
&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Every cell in the chunking column links to a source. A “yes” links to the feature's own documentation. An “app-side” links to that vendor's &lt;em&gt;own&lt;/em&gt; guidance on chunking in your application — which is what they publish instead of an in-engine option. If we missed a feature, or one has shipped since, &lt;a href="https://mnt.cr/go/pzyySW" rel="noopener noreferrer"&gt;tell us&lt;/a&gt; and we will fix it.&lt;/p&gt;

&lt;p&gt;Versions checked - 4 September 2026&lt;/p&gt;

&lt;p&gt;The latest stable release of each product available that day:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Product&lt;/th&gt;
&lt;th&gt;Version&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Elasticsearch&lt;/td&gt;
&lt;td&gt;9.5.3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenSearch&lt;/td&gt;
&lt;td&gt;3.8.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vespa&lt;/td&gt;
&lt;td&gt;8.750.13&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Azure AI Search&lt;/td&gt;
&lt;td&gt;REST API 2026-04-01&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PostgreSQL + pgai&lt;/td&gt;
&lt;td&gt;extension 0.11.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Milvus / Zilliz&lt;/td&gt;
&lt;td&gt;2.6.23&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qdrant&lt;/td&gt;
&lt;td&gt;1.19.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Weaviate&lt;/td&gt;
&lt;td&gt;1.38.13&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Meilisearch&lt;/td&gt;
&lt;td&gt;1.53.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typesense&lt;/td&gt;
&lt;td&gt;30.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Apache Solr&lt;/td&gt;
&lt;td&gt;10.0.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pinecone, MongoDB Atlas&lt;/td&gt;
&lt;td&gt;hosted services, no version to pin&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two things stand out.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Chunking in-engine is still rare.&lt;/strong&gt; Milvus, Qdrant, Weaviate, Pinecone, MongoDB Atlas, Typesense, Meilisearch and — since the 9.8 LLM module — Apache Solr will all run the embedding model for you, and every one of them will happily truncate your 4,000-word document without saying so. The splitting is your problem, in your application, in a language and a library that has no idea what tokenizer the model uses.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where chunking exists, the plumbing usually leaks.&lt;/strong&gt; OpenSearch gets you there with a &lt;code&gt;text_chunking&lt;/code&gt; processor feeding a &lt;code&gt;text_embedding&lt;/code&gt; processor writing into a nested field, queried with a nested query and a score mode. Azure AI Search wants a skillset with a Split skill, an embedding skill and index projections — and returns one result row per chunk, so grouping back to documents is on you. pgai Vectorizer writes chunks to a second table, so every query is a join plus a &lt;code&gt;DISTINCT ON&lt;/code&gt;. Elasticsearch's &lt;code&gt;semantic_text&lt;/code&gt; is genuinely close to Manticore's model: chunking settings on the inference endpoint, chunks hidden inside the field, one hit per document.&lt;/p&gt;

&lt;p&gt;Manticore does the same thing with less surface area: the strategy is an option on the column, the chunks are the column's value, and search returns documents. If you are weighing the whole stack rather than this one feature, we have written up the comparison with &lt;a href="https://mnt.cr/go/Xol6vq" rel="noopener noreferrer"&gt;Elasticsearch&lt;/a&gt; and with &lt;a href="https://mnt.cr/go/Od80kg" rel="noopener noreferrer"&gt;Turbopuffer&lt;/a&gt; too.&lt;/p&gt;

&lt;h2&gt;
  
  
  What chunking does not fix
&lt;/h2&gt;

&lt;p&gt;Chunking solves one problem well — a document longer than the model's window is no longer half-invisible. It does not make retrieval perfect, and two known gaps are worth naming.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A chunk does not know where it came from.&lt;/strong&gt; Split a document and you get a paragraph that says "do one node at a time and confirm every peer reports as synced" with no indication of &lt;em&gt;what&lt;/em&gt; is being rotated, or which product it belongs to. Anthropic's &lt;a href="https://mnt.cr/go/x9BIu6" rel="noopener noreferrer"&gt;contextual retrieval&lt;/a&gt; work put numbers on this: prepending a short, chunk-specific description of the surrounding document before embedding cut top-20 retrieval failures by 35%, and by 49% combined with a contextual BM25 index.&lt;/p&gt;

&lt;p&gt;Manticore does not do this for you. &lt;code&gt;FROM&lt;/code&gt; joins its fields with a space &lt;em&gt;before&lt;/em&gt; chunking, so listing &lt;code&gt;title&lt;/code&gt; first puts the title at the head of the text that gets split — which means it lands in the &lt;strong&gt;first&lt;/strong&gt; chunk and only that one. Every chunk after it is on its own:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- 'title' leads, so its words are in chunk 1; chunks 2..N never see them&lt;/span&gt;
&lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'title,body'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you need every chunk to carry context, you have to build it into the stored text yourself before inserting — for example by repeating a short heading at the start of each section of &lt;code&gt;body&lt;/code&gt;. There is no per-chunk prefix option today.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Chunk boundaries are decided before the model sees the text.&lt;/strong&gt; Manticore splits, then embeds each piece independently — the standard approach, and what every engine with in-engine chunking in the table above does. An alternative called &lt;a href="https://mnt.cr/go/PTrm05" rel="noopener noreferrer"&gt;late chunking&lt;/a&gt; inverts it: run a long-context model over the whole document first, then pool the token embeddings into chunks, so each chunk vector carries context from the rest of the document. It needs a long-context model and more compute per document, and Manticore does not do it today. If your documents depend heavily on cross-paragraph context, it is worth knowing the option exists.&lt;/p&gt;

&lt;p&gt;Neither gap changes the basic result: for long documents, chunked retrieval beats truncated retrieval by a wide margin, and reaching it means adding &lt;code&gt;chunk_strategy&lt;/code&gt; to one column.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits and gotchas
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;max_chunks&lt;/code&gt; discards text.&lt;/strong&gt; Manticore merges overflow into the last kept chunk, then truncates it to the model's window. Nothing warns you. It's a guard rail for outliers, not a way to save memory across the board.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Remote models chunk by bytes, not tokens.&lt;/strong&gt; OpenAI, Voyage and Jina have no local tokenizer, so Manticore falls back to a deliberately conservative estimate of &lt;strong&gt;3 bytes per token&lt;/strong&gt; — a chunk lands under the provider's cap rather than over it. In practice &lt;code&gt;max_tokens='N'&lt;/code&gt; becomes an &lt;code&gt;N × 3&lt;/code&gt;-byte window. We measured it against a stub endpoint with a 3,599-byte document and the &lt;code&gt;fixed&lt;/code&gt; strategy:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;code&gt;max_tokens&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;Byte window&lt;/th&gt;
&lt;th&gt;Chunks produced&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;300&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;td&gt;600&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;400&lt;/td&gt;
&lt;td&gt;1,200&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;English prose runs closer to 4 bytes per token, so on a remote model you get chunks roughly a quarter smaller than the number you asked for — set &lt;code&gt;max_tokens&lt;/code&gt; about 30% higher than you would for a local model to land in the same place. If exact boundaries matter, use a local model, where splitting is done on the model's real tokens.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Multi-vector columns can't be added with &lt;code&gt;ALTER&lt;/code&gt;.&lt;/strong&gt; &lt;code&gt;ALTER TABLE ... ADD COLUMN&lt;/code&gt; and &lt;code&gt;ALTER TABLE ... REBUILD EMBEDDINGS&lt;/code&gt; on a model-backed &lt;code&gt;float_vector_array&lt;/code&gt; are rejected. Recreate the table instead. Both work normally on a &lt;code&gt;float_vector&lt;/code&gt;, including with &lt;code&gt;mean&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Chunking applies to auto embeddings only.&lt;/strong&gt; Vectors you insert yourself are stored exactly as given — Manticore never re-cuts data you supplied. &lt;code&gt;chunk_strategy&lt;/code&gt; without &lt;code&gt;model_name&lt;/code&gt; is a DDL error, on purpose.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;embeddings&lt;/code&gt; is a reserved word.&lt;/strong&gt; &lt;code&gt;EMBEDDINGS&lt;/code&gt; is a DDL keyword (&lt;code&gt;ALTER TABLE ... REBUILD EMBEDDINGS&lt;/code&gt;), so a column literally named &lt;code&gt;embeddings&lt;/code&gt; is a syntax error unless escaped. Use escaping if you need that name.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Queries are not chunked.&lt;/strong&gt; A query is embedded whole, as a single vector. That is what you want: chunking exists to make a long &lt;em&gt;document&lt;/em&gt; findable, not to split a fifteen-word question.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The DDL tells you when a combination is wrong&lt;/strong&gt;, at &lt;code&gt;CREATE TABLE&lt;/code&gt; time rather than at the first insert:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;mysql&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;title&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="n"&gt;float_vector&lt;/span&gt; &lt;span class="p"&gt;...&lt;/span&gt; &lt;span class="n"&gt;chunk_strategy&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'sentence'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="n"&gt;ERROR&lt;/span&gt; &lt;span class="mi"&gt;1064&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;chunk_strategy&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'sentence'&lt;/span&gt; &lt;span class="n"&gt;produces&lt;/span&gt; &lt;span class="n"&gt;several&lt;/span&gt; &lt;span class="n"&gt;vectors&lt;/span&gt; &lt;span class="n"&gt;per&lt;/span&gt; &lt;span class="n"&gt;document&lt;/span&gt;
            &lt;span class="k"&gt;and&lt;/span&gt; &lt;span class="n"&gt;requires&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="n"&gt;float_vector_array&lt;/span&gt; &lt;span class="n"&gt;attribute&lt;/span&gt;

&lt;span class="n"&gt;mysql&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;...&lt;/span&gt; &lt;span class="n"&gt;chunk_strategy&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'fixed'&lt;/span&gt; &lt;span class="n"&gt;overlap_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'32'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="n"&gt;ERROR&lt;/span&gt; &lt;span class="mi"&gt;1064&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;overlap_tokens&lt;/span&gt; &lt;span class="n"&gt;requires&lt;/span&gt; &lt;span class="n"&gt;an&lt;/span&gt; &lt;span class="n"&gt;explicit&lt;/span&gt; &lt;span class="n"&gt;non&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;zero&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;

&lt;span class="n"&gt;mysql&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;...&lt;/span&gt; &lt;span class="n"&gt;chunk_strategy&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'paragraph'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="n"&gt;ERROR&lt;/span&gt; &lt;span class="mi"&gt;1064&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;unknown&lt;/span&gt; &lt;span class="n"&gt;chunk_strategy&lt;/span&gt; &lt;span class="s1"&gt;'paragraph'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="n"&gt;expected&lt;/span&gt; &lt;span class="k"&gt;truncate&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;mean&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;fixed&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="k"&gt;recursive&lt;/span&gt; &lt;span class="k"&gt;or&lt;/span&gt; &lt;span class="n"&gt;sentence&lt;/span&gt;

&lt;span class="n"&gt;mysql&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;...&lt;/span&gt; &lt;span class="n"&gt;chunk_strategy&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'truncate'&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'128'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="n"&gt;ERROR&lt;/span&gt; &lt;span class="mi"&gt;1064&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;chunk_strategy&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'truncate'&lt;/span&gt; &lt;span class="n"&gt;ignores&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;overlap_tokens&lt;/span&gt; &lt;span class="k"&gt;and&lt;/span&gt; &lt;span class="n"&gt;max_chunks&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;The shortest path to a working chunked semantic search:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;DROP&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;IF&lt;/span&gt; &lt;span class="k"&gt;EXISTS&lt;/span&gt; &lt;span class="n"&gt;docs&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;docs&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;title&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;content&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;chunks&lt;/span&gt; &lt;span class="n"&gt;float_vector_array&lt;/span&gt; &lt;span class="n"&gt;knn_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'hnsw'&lt;/span&gt; &lt;span class="n"&gt;hnsw_similarity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'cosine'&lt;/span&gt;
    &lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'Xenova/all-MiniLM-L6-v2'&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'title,content'&lt;/span&gt;
    &lt;span class="n"&gt;chunk_strategy&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'recursive'&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'256'&lt;/span&gt; &lt;span class="n"&gt;overlap_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'32'&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;INSERT&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;docs&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;VALUES&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'Backup and restore runbook'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'Nightly backups run at 02:00 UTC ... '&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;knn_dist&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;docs&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;knn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'how do I rotate the replication certificate'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No model to download by hand, no splitter to pick, no pipeline to maintain. One column option, and the parts of your documents that used to be invisible start showing up in results.&lt;/p&gt;

&lt;p&gt;Full reference: &lt;a href="https://mnt.cr/go/3aXp5K" rel="noopener noreferrer"&gt;Chunking strategies&lt;/a&gt; and &lt;a href="https://mnt.cr/go/7D6PF2" rel="noopener noreferrer"&gt;Multiple vectors per document&lt;/a&gt; in the manual. Questions and bug reports on &lt;a href="https://mnt.cr/go/pzyySW" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>database</category>
      <category>search</category>
      <category>ai</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Manticore Search 29.9.0: Chunked auto-embeddings and mmap columnar access</title>
      <dc:creator>Sergey Nikolaev</dc:creator>
      <pubDate>Fri, 11 Sep 2026 16:05:04 +0000</pubDate>
      <link>https://dev.to/sanikolaev/manticore-search-2990-chunked-auto-embeddings-and-mmap-columnar-access-28ic</link>
      <guid>https://dev.to/sanikolaev/manticore-search-2990-chunked-auto-embeddings-and-mmap-columnar-access-28ic</guid>
      <description>&lt;p&gt;Originally published on &lt;a href="https://mnt.cr/go/c72RPs" rel="noopener noreferrer"&gt;https://mnt.cr/go/c72RPs&lt;/a&gt; on September 11, 2026&lt;/p&gt;

&lt;h1&gt;
  
  
  Manticore Search 29.9.0: Chunked auto-embeddings and mmap columnar access
&lt;/h1&gt;

&lt;p&gt;Manticore Search 29.9.0 adds chunked and multi-vector auto-embeddings, embedding input limits, UTF-8 identifiers, mmap columnar access by default, and fixes for hybrid search, KNN, bulk ingestion, and grouping.&lt;/p&gt;

&lt;p&gt;[Manticore Search 29.9.0](&lt;a href="https://mnt.cr/go/LfXmU1" rel="noopener noreferrer"&gt;https://mnt.cr/go/LfXmU1&lt;/a&gt; has been released. The largest change is in auto-embeddings: long documents can now be split into searchable chunks, and a document can keep several vectors instead of compressing all of its content into one. This release also makes &lt;code&gt;mmap&lt;/code&gt; the default access mode for columnar attributes, adds safer controls for embedding workloads, UTF-8 identifiers, better backup support, and fixes across hybrid search, KNN, bulk ingestion, grouped search, and schema changes.&lt;/p&gt;

&lt;p&gt;This post covers everything shipped after 29.0.2, from &lt;strong&gt;29.0.3 through 29.9.0&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;❤️ We’d like to thank [@tudorvasinca](&lt;a href="https://mnt.cr/go/Vn5RXy" rel="noopener noreferrer"&gt;https://mnt.cr/go/Vn5RXy&lt;/a&gt; for their work on [PR #4857](&lt;a href="https://mnt.cr/go/DdBMb7" rel="noopener noreferrer"&gt;https://mnt.cr/go/DdBMb7&lt;/a&gt;, [PR #4859](&lt;a href="https://mnt.cr/go/jZePzd" rel="noopener noreferrer"&gt;https://mnt.cr/go/jZePzd&lt;/a&gt;, and [PR #4873](&lt;a href="https://mnt.cr/go/eQNGL6" rel="noopener noreferrer"&gt;https://mnt.cr/go/eQNGL6&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Upgrade notes
&lt;/h2&gt;

&lt;p&gt;There is no release-wide mandatory data migration. Existing tables and configuration can be upgraded normally, but a few fixes need follow-up if the earlier behavior already affected your data:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;German sharp-s normalization is opt-in. Switching an existing table to &lt;code&gt;lemmatize_de_v2&lt;/code&gt; or &lt;code&gt;lemmatize_de_v2_all&lt;/code&gt; changes indexed terms, so rebuild a plain table or replay documents into a new RT table.&lt;/li&gt;
&lt;li&gt;Authenticated &lt;code&gt;BACKUP&lt;/code&gt; introduces the &lt;code&gt;backup&lt;/code&gt; authorization action. Before downgrading to 29.3.12 or earlier, remove any &lt;code&gt;backup&lt;/code&gt; grants.&lt;/li&gt;
&lt;li&gt;Columnar attributes now use &lt;code&gt;mmap&lt;/code&gt; by default instead of buffered &lt;code&gt;file&lt;/code&gt; reads. Set &lt;code&gt;access_columnar_attrs='file'&lt;/code&gt; explicitly if you need to retain the previous access mode.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Long documents can keep more than one embedding
&lt;/h2&gt;

&lt;p&gt;The old auto-embedding path produced one vector per document and truncated text that did not fit the model input window. That works for titles and short descriptions, but it means a relevant paragraph near the end of a long article may never reach the index.&lt;/p&gt;

&lt;p&gt;Manticore Search now supports five [chunking strategies](&lt;a href="https://mnt.cr/go/BXIJ8c" rel="noopener noreferrer"&gt;https://mnt.cr/go/BXIJ8c&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;truncate&lt;/code&gt; keeps the previous behavior and remains the default.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;mean&lt;/code&gt; embeds every chunk and averages the results into one vector.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;fixed&lt;/code&gt; splits text into fixed-size token windows.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;recursive&lt;/code&gt; prefers paragraph, line, sentence, and space boundaries in that order.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;sentence&lt;/code&gt; groups complete sentences up to the configured limit.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The multi-vector strategies use [&lt;code&gt;float_vector_array&lt;/code&gt;](&lt;a href="https://mnt.cr/go/Y31JKu" rel="noopener noreferrer"&gt;https://mnt.cr/go/Y31JKu&lt;/a&gt; Each chunk competes independently during KNN search, but Manticore returns the document once and uses its closest chunk for &lt;code&gt;knn_dist()&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;articles&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;title&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;content&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;chunks&lt;/span&gt; &lt;span class="n"&gt;float_vector_array&lt;/span&gt; &lt;span class="n"&gt;knn_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'hnsw'&lt;/span&gt; &lt;span class="n"&gt;hnsw_similarity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'cosine'&lt;/span&gt;
    &lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'Xenova/all-MiniLM-L6-v2'&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'title,content'&lt;/span&gt;
    &lt;span class="n"&gt;chunk_strategy&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'sentence'&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'256'&lt;/span&gt; &lt;span class="n"&gt;overlap_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'32'&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;INSERT&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;articles&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;VALUES&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'Rotating certificates'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'A long guide with many sections ...'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;knn_dist&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;articles&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;knn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'how do I rotate a certificate'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;MAX_TOKENS&lt;/code&gt;, &lt;code&gt;OVERLAP_TOKENS&lt;/code&gt;, and &lt;code&gt;MAX_CHUNKS&lt;/code&gt; control chunk size, shared context at boundaries, and the maximum vector count. Declare a model-backed &lt;code&gt;float_vector_array&lt;/code&gt; when creating the table: adding one later with &lt;code&gt;ALTER TABLE ... ADD COLUMN&lt;/code&gt; and rebuilding its embeddings are not supported yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  Put a ceiling on local embedding work
&lt;/h2&gt;

&lt;p&gt;Long-context models can make a single large input unexpectedly expensive, especially on CPU. The new [&lt;code&gt;MAX_INPUT_TOKENS&lt;/code&gt;](&lt;a href="https://mnt.cr/go/sX4dok" rel="noopener noreferrer"&gt;https://mnt.cr/go/sX4dok&lt;/a&gt; column option caps how much of each input is sent to a local embedding model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;articles&lt;/span&gt;
&lt;span class="k"&gt;MODIFY&lt;/span&gt; &lt;span class="k"&gt;COLUMN&lt;/span&gt; &lt;span class="n"&gt;chunks&lt;/span&gt; &lt;span class="n"&gt;MAX_INPUT_TOKENS&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'512'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The change applies to embeddings generated afterward; existing vectors stay as they are. Set it during &lt;code&gt;CREATE TABLE&lt;/code&gt; or change it later without re-embedding the table. A value of &lt;code&gt;0&lt;/code&gt;, or leaving the option out, uses the model's own limit.&lt;/p&gt;

&lt;p&gt;This release also fixes two less visible problems around this path. Models no longer remain cached after an embedding column is modified, and configurations with different &lt;code&gt;API_TIMEOUT&lt;/code&gt; or &lt;code&gt;MAX_INPUT_TOKENS&lt;/code&gt; values no longer collide and reuse the wrong cached model.&lt;/p&gt;




&lt;h2&gt;
  
  
  UTF-8 identifiers and better German matching
&lt;/h2&gt;

&lt;p&gt;Table, field, and attribute names now follow one consistent [safe UTF-8 identifier syntax](&lt;a href="https://mnt.cr/go/mfKS2T" rel="noopener noreferrer"&gt;https://mnt.cr/go/mfKS2T&lt;/a&gt; RT, percolate, distributed, template, and plain tables can use localized names such as Chinese or Cyrillic identifiers across DDL, expressions, field selectors, and inferred source schemas.&lt;/p&gt;

&lt;p&gt;German AOT morphology also gains opt-in sharp-s normalization. With &lt;code&gt;charset_table=non_cont,german&lt;/code&gt; and &lt;code&gt;morphology=lemmatize_de_v2&lt;/code&gt; or &lt;code&gt;lemmatize_de_v2_all&lt;/code&gt;, forms such as &lt;code&gt;Straße&lt;/code&gt;, &lt;code&gt;Strasse&lt;/code&gt;, and &lt;code&gt;STRAẞE&lt;/code&gt; match in ordinary whole-word searches. If &lt;code&gt;index_exact_words=1&lt;/code&gt; is enabled, exact-word queries can still distinguish the &lt;code&gt;ß&lt;/code&gt; and &lt;code&gt;ss&lt;/code&gt; forms.&lt;/p&gt;




&lt;h2&gt;
  
  
  Better load testing and backups
&lt;/h2&gt;

&lt;p&gt;[&lt;code&gt;manticore-load&lt;/code&gt;](&lt;a href="https://mnt.cr/go/cnoLS6" rel="noopener noreferrer"&gt;https://mnt.cr/go/cnoLS6&lt;/a&gt; can now benchmark through the HTTP JSON API with &lt;code&gt;--http&lt;/code&gt;. Writes use &lt;code&gt;/bulk&lt;/code&gt;, searches use &lt;code&gt;/search&lt;/code&gt;, and &lt;code&gt;--table&lt;/code&gt; selects the target table. Its reports now include local &lt;code&gt;searchd&lt;/code&gt; RSS during the run plus peak RSS, disk, and CPU statistics at the end, including aggregate monitoring for multi-command workloads.&lt;/p&gt;

&lt;p&gt;[Manticore Backup](&lt;a href="https://mnt.cr/go/t4acW4" rel="noopener noreferrer"&gt;https://mnt.cr/go/t4acW4&lt;/a&gt; now works with authenticated Manticore Search installations through username/password or bearer-token credentials. SQL [&lt;code&gt;BACKUP&lt;/code&gt;](&lt;a href="https://mnt.cr/go/6dxdrT" rel="noopener noreferrer"&gt;https://mnt.cr/go/6dxdrT&lt;/a&gt; has a dedicated authorization action and checks read access to the selected tables.&lt;/p&gt;

&lt;p&gt;S3 backups and restores can also use the AWS SDK credential provider chain when static keys are not set. That includes IRSA, shared credentials, ECS task roles, and EC2 instance profiles. Temporary credentials can supply &lt;code&gt;AWS_SESSION_TOKEN&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Columnar attributes use mmap by default
&lt;/h2&gt;

&lt;p&gt;Manticore Search now defaults [&lt;code&gt;access_columnar_attrs&lt;/code&gt;](&lt;a href="https://mnt.cr/go/sdnJ5l" rel="noopener noreferrer"&gt;https://mnt.cr/go/sdnJ5l&lt;/a&gt; to &lt;code&gt;mmap&lt;/code&gt;. The operating system maps and caches &lt;code&gt;*.spc&lt;/code&gt; columnar-attribute files on demand, without prereading the whole file at startup. The previous buffered path remains available with &lt;code&gt;access_columnar_attrs='file'&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;ALTER TABLE&lt;/code&gt; also reopens replaced columnar storage with the configured access mode, so an altered table no longer falls back to a different reader than the one requested.&lt;/p&gt;




&lt;h2&gt;
  
  
  Vector and grouped-search fixes
&lt;/h2&gt;

&lt;p&gt;Several fixes target queries that were valid but could return incomplete results or fail under a particular table layout:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Distributed and sharded KNN queries with a local shard no longer rescore merged 1-bit-quantized results twice, which could crash the coordinator or return the wrong nearest neighbor. ([Issue #4791](&lt;a href="https://mnt.cr/go/LfXW1t" rel="noopener noreferrer"&gt;https://mnt.cr/go/LfXW1t&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;KNN queries with additional filters avoid a redundant &lt;code&gt;knn_dist&lt;/code&gt; prefilter when HNSW already excludes documents without vectors. ([PR #4861](&lt;a href="https://mnt.cr/go/bZtpIJ" rel="noopener noreferrer"&gt;https://mnt.cr/go/bZtpIJ&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;LENGTH()&lt;/code&gt; on a &lt;code&gt;float_vector_array&lt;/code&gt; now reports the number of vectors rather than its internal storage-word count. ([PR #4879](&lt;a href="https://mnt.cr/go/8be2oV" rel="noopener noreferrer"&gt;https://mnt.cr/go/8be2oV&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Hybrid search with &lt;code&gt;GROUP BY&lt;/code&gt; retains all buckets, including MVA groups, and follows both final and within-group ordering. ([Issue #4639](&lt;a href="https://mnt.cr/go/b5Hmg8" rel="noopener noreferrer"&gt;https://mnt.cr/go/b5Hmg8&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Hybrid filters on &lt;code&gt;weight()&lt;/code&gt; and expressions or aliases derived from it now run after fusion against the final text weight instead of being ignored. Weight-dependent filters inside &lt;code&gt;OR&lt;/code&gt; trees remain unsupported and return an explicit error. ([Issue #4889](&lt;a href="https://mnt.cr/go/OMu9qt" rel="noopener noreferrer"&gt;https://mnt.cr/go/OMu9qt&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Multi-chunk RT grouping no longer risks duplicate groups, incorrect split counts, or a hang while ordering &lt;code&gt;COUNT(DISTINCT ...)&lt;/code&gt; results. ([Issue #4856](&lt;a href="https://mnt.cr/go/KtrFXK" rel="noopener noreferrer"&gt;https://mnt.cr/go/KtrFXK&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Document-ID filters whose signed representation is negative now follow the lookup index's unsigned ordering. ([Issue #4774](&lt;a href="https://mnt.cr/go/nEO1vc" rel="noopener noreferrer"&gt;https://mnt.cr/go/nEO1vc&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There are crash fixes here too: a second hybrid-search statement in a multi-statement request ([PR #4864](&lt;a href="https://mnt.cr/go/RJ5hGi" rel="noopener noreferrer"&gt;https://mnt.cr/go/RJ5hGi&lt;/a&gt;, distributed JSON aggregation sorted by a string attribute ([Issue #4822](&lt;a href="https://mnt.cr/go/Jdpl0M" rel="noopener noreferrer"&gt;https://mnt.cr/go/Jdpl0M&lt;/a&gt;, and dropping a table during auto-embedding precommit ([Issue #4860](&lt;a href="https://mnt.cr/go/Edp7gc" rel="noopener noreferrer"&gt;https://mnt.cr/go/Edp7gc&lt;/a&gt; are all handled safely now.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bulk ingestion behaves predictably
&lt;/h2&gt;

&lt;p&gt;Elasticsearch-compatible &lt;code&gt;/_bulk&lt;/code&gt; requests now return HTTP &lt;code&gt;200&lt;/code&gt; once a batch has been processed, while individual failures remain visible through &lt;code&gt;errors: true&lt;/code&gt; and per-item statuses. Duplicate &lt;code&gt;create&lt;/code&gt; actions return per-item &lt;code&gt;409&lt;/code&gt; &lt;code&gt;version_conflict_engine_exception&lt;/code&gt; errors, including duplicates within the same batch. This prevents clients such as Fluent Bit from retrying writes that already succeeded.&lt;/p&gt;

&lt;p&gt;Fixed-length gzip-compressed &lt;code&gt;/bulk&lt;/code&gt; bodies are also decoded correctly when they arrive across multiple socket reads. And if native bulk processing fails because the target table does not exist, the request can again reach Manticore's auto-schema fallback with valid NDJSON while preserving the expected bulk response envelope.&lt;/p&gt;




&lt;h2&gt;
  
  
  More reliability fixes
&lt;/h2&gt;

&lt;p&gt;The rest of the release closes a broad set of operational and compatibility problems:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;searchd --stopwait&lt;/code&gt; no longer hangs while a sharded table is being rebalanced after a node rejoins. ([Issue #3905](&lt;a href="https://mnt.cr/go/ZVFkNe" rel="noopener noreferrer"&gt;https://mnt.cr/go/ZVFkNe&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Compatible older binlogs replay safely during upgrade and completed RT chunks are published before clean shutdown. ([Issue #4808](&lt;a href="https://mnt.cr/go/O0ZRHJ" rel="noopener noreferrer"&gt;https://mnt.cr/go/O0ZRHJ&lt;/a&gt; Fatal replay diagnostics also name the relevant recovery flag. ([Issue #4811](&lt;a href="https://mnt.cr/go/pzGwEm" rel="noopener noreferrer"&gt;https://mnt.cr/go/pzGwEm&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;indexer&lt;/code&gt; creates missing parent directories for plain-table paths when the nearest existing parent is writable. ([Issue #4793](&lt;a href="https://mnt.cr/go/eIFSQf" rel="noopener noreferrer"&gt;https://mnt.cr/go/eIFSQf&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;UUID document IDs no longer cause stored text fields to come back empty. ([Issue #4833](&lt;a href="https://mnt.cr/go/otelYC" rel="noopener noreferrer"&gt;https://mnt.cr/go/otelYC&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;ALTER TABLE ... RENAME&lt;/code&gt; preserves hidden remote embedding API keys without exposing them in &lt;code&gt;SHOW CREATE TABLE&lt;/code&gt;. ([Issue #4842](&lt;a href="https://mnt.cr/go/XkT4iw" rel="noopener noreferrer"&gt;https://mnt.cr/go/XkT4iw&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Sequel Ace 5.3.1+ compatibility probes work again. ([Issue #4828](&lt;a href="https://mnt.cr/go/zjvXbT" rel="noopener noreferrer"&gt;https://mnt.cr/go/zjvXbT&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;JSON &lt;code&gt;/search&lt;/code&gt; keeps distances for negated &lt;code&gt;NEAR&lt;/code&gt; and proximity operators. ([Issue #4784](&lt;a href="https://mnt.cr/go/NgB0Q7" rel="noopener noreferrer"&gt;https://mnt.cr/go/NgB0Q7&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Internal string-sort helper columns no longer leak from &lt;code&gt;LEFT JOIN&lt;/code&gt; results. ([Issue #4788](&lt;a href="https://mnt.cr/go/rxNeFg" rel="noopener noreferrer"&gt;https://mnt.cr/go/rxNeFg&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Malformed binary API &lt;code&gt;SEARCH&lt;/code&gt; element counts are rejected instead of terminating &lt;code&gt;searchd&lt;/code&gt;. ([PR #4790](&lt;a href="https://mnt.cr/go/ufPVLK" rel="noopener noreferrer"&gt;https://mnt.cr/go/ufPVLK&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For the complete list, see the [Version 29.9.0 changelog](&lt;a href="https://mnt.cr/go/gl5v6N" rel="noopener noreferrer"&gt;https://mnt.cr/go/gl5v6N&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Get Manticore Search 29.9.0
&lt;/h2&gt;

&lt;p&gt;Install or upgrade Manticore Search with the [installation guide](&lt;a href="https://mnt.cr/go/XPyNrv" rel="noopener noreferrer"&gt;https://mnt.cr/go/XPyNrv&lt;/a&gt; Review the upgrade notes above if you use columnar attributes, German AOT morphology, or authenticated backups.&lt;/p&gt;

&lt;h2&gt;
  
  
  Need help or want to connect?
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Join our [Slack](&lt;a href="https://mnt.cr/go/v4HPeK" rel="noopener noreferrer"&gt;https://mnt.cr/go/v4HPeK&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Visit the [Forum](&lt;a href="https://mnt.cr/go/R7u2vk" rel="noopener noreferrer"&gt;https://mnt.cr/go/R7u2vk&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Report issues or suggest features on [GitHub](&lt;a href="https://mnt.cr/go/CURvTu" rel="noopener noreferrer"&gt;https://mnt.cr/go/CURvTu&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Email us at &lt;code&gt;contact@manticoresearch.com&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>database</category>
      <category>search</category>
      <category>vectordatabase</category>
    </item>
    <item>
      <title>Explore Manticore Data with OpenSearch Dashboards</title>
      <dc:creator>Sergey Nikolaev</dc:creator>
      <pubDate>Tue, 08 Sep 2026 10:00:00 +0000</pubDate>
      <link>https://dev.to/sanikolaev/explore-manticore-data-with-opensearch-dashboards-59db</link>
      <guid>https://dev.to/sanikolaev/explore-manticore-data-with-opensearch-dashboards-59db</guid>
      <description>&lt;p&gt;Originally published on &lt;a href="https://manticoresearch.com/blog/manticore-opensearch-dashboards-integration/" rel="noopener noreferrer"&gt;https://manticoresearch.com/blog/manticore-opensearch-dashboards-integration/&lt;/a&gt; on September 8, 2026&lt;/p&gt;

&lt;h1&gt;
  
  
  Explore Manticore Data with OpenSearch Dashboards
&lt;/h1&gt;

&lt;p&gt;Connect OpenSearch Dashboards to Manticore Search, explore data in Discover, and build visualizations and dashboards using a familiar interface.&lt;/p&gt;

&lt;p&gt;OpenSearch Dashboards provides a familiar visual interface for exploring data, building charts, and assembling interactive dashboards. Although it is normally paired with OpenSearch, you can also connect it directly to Manticore Search.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7wd6g799m5lqto6r8kv1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7wd6g799m5lqto6r8kv1.png" alt="Main view" width="800" height="417"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The integration lets you keep Manticore as the search and analytics backend while using &lt;strong&gt;Discover&lt;/strong&gt;, &lt;strong&gt;Visualize&lt;/strong&gt;, and &lt;strong&gt;Dashboards&lt;/strong&gt; as the user interface. This is especially useful for log and event data: tools such as Logstash, Filebeat, Fluent Bit, and Vector can send data to Manticore, while OpenSearch Dashboards gives your team a convenient way to inspect it.&lt;/p&gt;

&lt;p&gt;In this tutorial, we will configure the connection, add a small sample dataset, explore it in Discover, and create a visualization.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the integration works
&lt;/h2&gt;

&lt;p&gt;OpenSearch Dashboards communicates with its backend through an HTTP API. Manticore exposes a compatible API on its HTTP listener, which uses port &lt;code&gt;9308&lt;/code&gt; by default.&lt;/p&gt;

&lt;p&gt;The compatibility layer is provided by the &lt;strong&gt;EmulateElastic&lt;/strong&gt; plugin in &lt;a href="https://manual.manticoresearch.com/Installation/Manticore_Buddy" rel="noopener noreferrer"&gt;Manticore Buddy&lt;/a&gt;. It handles the requests that OpenSearch Dashboards needs for startup, index discovery, searches, filters, and supported aggregations. Buddy normally starts automatically with &lt;code&gt;searchd&lt;/code&gt;, so a standard Manticore installation already has the required component.&lt;/p&gt;

&lt;p&gt;This is an API compatibility integration, not an embedded OpenSearch cluster. Manticore stores and queries the data, and OpenSearch Dashboards provides the visual interface.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;

&lt;p&gt;For this walkthrough, you need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A running Manticore Search instance with its HTTP endpoint available at &lt;code&gt;http://localhost:9308&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Manticore Buddy installed and running&lt;/li&gt;
&lt;li&gt;OpenSearch Dashboards &lt;strong&gt;3.4.0&lt;/strong&gt;, the currently tested and recommended version&lt;/li&gt;
&lt;li&gt;Manticore running in real-time mode&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Other OpenSearch Dashboards versions may work, but they have not been tested to the same extent. The version reported by Manticore must also match the version of OpenSearch Dashboards.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: Configure Manticore
&lt;/h2&gt;

&lt;p&gt;Open your Manticore configuration file and make sure that the HTTP listener is enabled. Set &lt;code&gt;kibana_version_string&lt;/code&gt; to the version of OpenSearch Dashboards that you plan to run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ini"&gt;&lt;code&gt;&lt;span class="err"&gt;searchd&lt;/span&gt; &lt;span class="err"&gt;{&lt;/span&gt;
    &lt;span class="py"&gt;listen&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;127.0.0.1:9308:http&lt;/span&gt;
    &lt;span class="py"&gt;pid_file&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;/var/run/manticore/searchd.pid&lt;/span&gt;
    &lt;span class="py"&gt;data_dir&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;/var/lib/manticore&lt;/span&gt;
    &lt;span class="py"&gt;kibana_version_string&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;3.4.0&lt;/span&gt;
&lt;span class="err"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The name &lt;code&gt;kibana_version_string&lt;/code&gt; is retained for compatibility with the existing Elasticsearch-style API. OpenSearch Dashboards checks the backend version during startup, so a mismatch can produce warnings or prevent the application from starting.&lt;/p&gt;

&lt;p&gt;Restart Manticore after changing the configuration.&lt;/p&gt;

&lt;p&gt;If Manticore and OpenSearch Dashboards run on different hosts or in different containers, do not bind the HTTP listener only to &lt;code&gt;127.0.0.1&lt;/code&gt;. Bind it to an address that OpenSearch Dashboards can reach, and restrict access with your network or firewall configuration.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: Configure OpenSearch Dashboards
&lt;/h2&gt;

&lt;p&gt;Open &lt;code&gt;opensearch_dashboards.yml&lt;/code&gt;. In a &lt;a href="https://docs.opensearch.org/latest/install-and-configure/install-opensearch/tar/" rel="noopener noreferrer"&gt;tarball installation&lt;/a&gt;, it is usually located at &lt;code&gt;config/opensearch_dashboards.yml&lt;/code&gt;; packages may place it at &lt;code&gt;/etc/opensearch-dashboards/opensearch_dashboards.yml&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Point &lt;code&gt;opensearch.hosts&lt;/code&gt; to Manticore's HTTP endpoint:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;opensearch.hosts&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;http://localhost:9308"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Manticore does not provide the OpenSearch Security plugin. The corresponding Dashboards plugin must therefore be disabled.&lt;/p&gt;

&lt;p&gt;For a tarball installation, stop OpenSearch Dashboards and remove the plugin:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;./bin/opensearch-dashboards-plugin remove securityDashboards
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you run OpenSearch Dashboards in Docker, provide both settings through environment variables:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;environment&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;OPENSEARCH_HOSTS&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;["http://manticore:9308"]'&lt;/span&gt;
  &lt;span class="na"&gt;DISABLE_SECURITY_DASHBOARDS_PLUGIN&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;true"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here, &lt;code&gt;manticore&lt;/code&gt; is the hostname or container name reachable from the OpenSearch Dashboards container. Using &lt;code&gt;localhost&lt;/code&gt; inside that container would refer to the container itself, not to Manticore.&lt;/p&gt;

&lt;p&gt;Start OpenSearch Dashboards and open &lt;a href="http://localhost:5601" rel="noopener noreferrer"&gt;http://localhost:5601&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: Add sample data to Manticore
&lt;/h2&gt;

&lt;p&gt;If you already have a real-time table, you can skip this step. Otherwise, connect to Manticore through its MySQL-compatible port:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;mysql &lt;span class="nt"&gt;-h127&lt;/span&gt;.0.0.1 &lt;span class="nt"&gt;-P9306&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Create a table for a small set of application events:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;app_events&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;message&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;service&lt;/span&gt; &lt;span class="n"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;status_code&lt;/span&gt; &lt;span class="n"&gt;uint&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;response_time&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;event_time&lt;/span&gt; &lt;span class="nb"&gt;timestamp&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Add a few documents:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;INSERT&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;app_events&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;status_code&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;response_time&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;event_time&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;VALUES&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'Request completed'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'catalog'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;12&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1786924800&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'Request completed'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'checkout'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;31&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1786928400&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'Upstream timeout'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'checkout'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;504&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;75&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1786932000&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'Product not found'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'catalog'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;404&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;08&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1786935600&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'Request completed'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'catalog'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;15&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1786939200&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'Payment rejected'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'payments'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;422&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;44&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1786942800&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This schema gives us text to search, dimensions to group by, metrics to aggregate, and a timestamp for time-based charts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: Create an index pattern
&lt;/h2&gt;

&lt;p&gt;OpenSearch Dashboards uses an &lt;strong&gt;index pattern&lt;/strong&gt; to select the data source shown in Discover and visualizations. In this integration, the pattern matches a Manticore table name.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Open &lt;strong&gt;Management &amp;gt; Dashboards Management&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Go to &lt;strong&gt;Index Patterns&lt;/strong&gt; and choose &lt;strong&gt;Create index pattern&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Enter &lt;code&gt;app_events&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Select &lt;code&gt;event_time&lt;/code&gt; as the time field.&lt;/li&gt;
&lt;li&gt;Save the index pattern.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The &lt;code&gt;app_events&lt;/code&gt; fields should now be available in OpenSearch Dashboards.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 5: Explore the data in Discover
&lt;/h2&gt;

&lt;p&gt;Open &lt;strong&gt;Discover&lt;/strong&gt; and select the &lt;code&gt;app_events&lt;/code&gt; index pattern. You can now:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Inspect individual documents&lt;/li&gt;
&lt;li&gt;Search the &lt;code&gt;message&lt;/code&gt; field&lt;/li&gt;
&lt;li&gt;Filter by fields such as &lt;code&gt;service&lt;/code&gt; or &lt;code&gt;status_code&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Change the time range&lt;/li&gt;
&lt;li&gt;Add and remove columns from the results&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For example, add a filter where &lt;code&gt;status_code&lt;/code&gt; is greater than or equal to &lt;code&gt;400&lt;/code&gt; to focus on failed requests. You can then add &lt;code&gt;service&lt;/code&gt;, &lt;code&gt;message&lt;/code&gt;, &lt;code&gt;status_code&lt;/code&gt;, and &lt;code&gt;response_time&lt;/code&gt; as columns to get a compact error view.&lt;/p&gt;

&lt;p&gt;Simple Dashboard Query Language searches work with the integration. Advanced DQL features—including nested field searches, regular expressions, fuzzy searches, proximity searches, and term boosting—may not be compatible with Manticore.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 6: Build a visualization
&lt;/h2&gt;

&lt;p&gt;Open &lt;strong&gt;Visualize&lt;/strong&gt;, create a new visualization, and choose the &lt;code&gt;app_events&lt;/code&gt; index pattern. OpenSearch Dashboards can use the following bucket aggregations with Manticore:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;terms&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;histogram&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;date_histogram&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;range&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;date_range&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Supported metric aggregations include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;max&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;min&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;sum&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;avg&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;As a first chart, create a bar chart of events by service:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Use &lt;code&gt;Count&lt;/code&gt; as the metric.&lt;/li&gt;
&lt;li&gt;Add a &lt;code&gt;Terms&lt;/code&gt; bucket using the &lt;code&gt;service&lt;/code&gt; field.&lt;/li&gt;
&lt;li&gt;Apply the changes.&lt;/li&gt;
&lt;li&gt;Save the visualization as &lt;strong&gt;Events by service&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;You can also create a time-series chart with a &lt;code&gt;Date Histogram&lt;/code&gt; on &lt;code&gt;event_time&lt;/code&gt;, or chart average latency by using the &lt;code&gt;avg&lt;/code&gt; metric on &lt;code&gt;response_time&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;After saving your visualizations, open &lt;strong&gt;Dashboards&lt;/strong&gt;, create a dashboard, and add them. Filters applied at dashboard level let you focus the whole view on one service, status code, or time range.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is supported
&lt;/h2&gt;

&lt;p&gt;The integration covers the main data exploration workflow:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Searching and filtering documents in Discover&lt;/li&gt;
&lt;li&gt;Creating index patterns for Manticore tables&lt;/li&gt;
&lt;li&gt;Building visualizations with supported bucket and metric aggregations&lt;/li&gt;
&lt;li&gt;Saving visualizations and combining them into dashboards&lt;/li&gt;
&lt;li&gt;Managing index patterns and saved objects in Dashboards Management&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Manticore also emulates the stack-level requests needed during OpenSearch Dashboards startup, including node version information, cluster settings, configuration objects, and index listings.&lt;/p&gt;

&lt;h2&gt;
  
  
  Current limitations
&lt;/h2&gt;

&lt;p&gt;OpenSearch Dashboards includes many features that depend on OpenSearch-specific APIs or field types. Those features are outside the scope of this integration.&lt;/p&gt;

&lt;p&gt;Unsupported field types include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Geographic and Cartesian fields such as &lt;code&gt;geo_point&lt;/code&gt;, &lt;code&gt;geo_shape&lt;/code&gt;, &lt;code&gt;xy_point&lt;/code&gt;, and &lt;code&gt;xy_shape&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Range types such as &lt;code&gt;integer_range&lt;/code&gt;, &lt;code&gt;ip_range&lt;/code&gt;, and &lt;code&gt;date_range&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Specialized search types such as &lt;code&gt;semantic&lt;/code&gt;, &lt;code&gt;rank_feature&lt;/code&gt;, and &lt;code&gt;percolator&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;OpenSearch vector types such as &lt;code&gt;knn_vector&lt;/code&gt; and &lt;code&gt;sparse_vector&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Relational types such as &lt;code&gt;nested&lt;/code&gt; and &lt;code&gt;join&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Advanced string types such as &lt;code&gt;completion&lt;/code&gt; and &lt;code&gt;search_as_you_type&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Plain &lt;code&gt;text&lt;/code&gt; and &lt;code&gt;keyword&lt;/code&gt; fields are supported. Manticore's own vector search remains available through its SQL and JSON APIs, but OpenSearch Dashboards cannot represent it through OpenSearch's vector field type.&lt;/p&gt;

&lt;p&gt;Nested aggregations—an &lt;code&gt;aggs&lt;/code&gt; block inside another &lt;code&gt;aggs&lt;/code&gt; block—are not supported. Metric functions are limited to those implemented by Manticore.&lt;/p&gt;

&lt;p&gt;OpenSearch-specific applications and administration tools are also not available, including:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Geospatial visualizations&lt;/li&gt;
&lt;li&gt;Observability and trace analytics&lt;/li&gt;
&lt;li&gt;Alerting and Anomaly Detection&lt;/li&gt;
&lt;li&gt;Security Analytics&lt;/li&gt;
&lt;li&gt;Index State Management and Index Management&lt;/li&gt;
&lt;li&gt;Performance Analyzer&lt;/li&gt;
&lt;li&gt;OpenSearch Security plugin workflows&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These limitations do not affect the core Discover, visualization, and dashboard workflow described above.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bringing in real data
&lt;/h2&gt;

&lt;p&gt;Once the connection works, you can replace the sample table with your own data pipeline. Manticore integrates with &lt;a href="https://manual.manticoresearch.com/Integration/Logstash" rel="noopener noreferrer"&gt;Logstash&lt;/a&gt;, &lt;a href="https://manual.manticoresearch.com/Integration/Filebeat" rel="noopener noreferrer"&gt;Filebeat&lt;/a&gt;, &lt;a href="https://manual.manticoresearch.com/Integration/Fluent_Bit" rel="noopener noreferrer"&gt;Fluent Bit&lt;/a&gt;, and &lt;a href="https://manual.manticoresearch.com/Integration/Vector" rel="noopener noreferrer"&gt;Vector&lt;/a&gt;. These tools can collect and transform logs or events before sending them to Manticore's Elasticsearch-compatible HTTP endpoint.&lt;/p&gt;

&lt;p&gt;The resulting workflow is straightforward:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;An agent or pipeline collects your data.&lt;/li&gt;
&lt;li&gt;Manticore indexes it in real time.&lt;/li&gt;
&lt;li&gt;OpenSearch Dashboards queries Manticore over HTTP.&lt;/li&gt;
&lt;li&gt;Users explore the data and build dashboards in a familiar interface.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The OpenSearch Dashboards integration gives Manticore users a practical visual layer for log analysis and data exploration. With a small amount of configuration, you can search documents in Discover, build charts from supported aggregations, and combine them into dashboards while Manticore handles storage and query execution.&lt;/p&gt;

&lt;p&gt;For the latest compatibility details and configuration notes, see the &lt;a href="https://manual.manticoresearch.com/dev/Integration/Opensearch_Dashboards" rel="noopener noreferrer"&gt;OpenSearch Dashboards integration documentation&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>database</category>
      <category>search</category>
      <category>tutorial</category>
      <category>devops</category>
    </item>
    <item>
      <title>How to return top-N from each group with GROUP_CONCAT()</title>
      <dc:creator>Sergey Nikolaev</dc:creator>
      <pubDate>Fri, 04 Sep 2026 10:00:05 +0000</pubDate>
      <link>https://dev.to/sanikolaev/how-to-return-top-n-from-each-group-with-groupconcat-1i3o</link>
      <guid>https://dev.to/sanikolaev/how-to-return-top-n-from-each-group-with-groupconcat-1i3o</guid>
      <description>&lt;p&gt;Originally posted by Stanislav Klinov on &lt;a href="https://manticoresearch.com/blog/group-concat-top-n-per-group/" rel="noopener noreferrer"&gt;https://manticoresearch.com/blog/group-concat-top-n-per-group/&lt;/a&gt; on September 4, 2026&lt;/p&gt;

&lt;h1&gt;
  
  
  How to return top-N from each group with GROUP_CONCAT()
&lt;/h1&gt;

&lt;p&gt;A practical example of how to select several most recent items from each group in Manticore Search and combine them into a single string.&lt;/p&gt;

&lt;p&gt;Imagine a customer support screen. An agent searches for refund-related events and, instead of a long log, wants a short summary: one row per user, the total number of matches, and the &lt;strong&gt;five&lt;/strong&gt; most recent events. If more details are needed, the application can load them by ID.&lt;/p&gt;

&lt;p&gt;Two obvious approaches don't quite give us what we need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;regular &lt;code&gt;GROUP_CONCAT()&lt;/code&gt; collects all values in the group together,&lt;/li&gt;
&lt;li&gt;while &lt;code&gt;GROUP N BY&lt;/code&gt; returns several rows per user.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Starting with &lt;a href="https://manticoresearch.com/blog/manticore-search-28-6-6/" rel="noopener noreferrer"&gt;Manticore Search 28.6.6&lt;/a&gt;, values inside &lt;code&gt;GROUP_CONCAT()&lt;/code&gt; can be sorted and limited to the number you need:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;GROUP_CONCAT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;event_ts&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt; &lt;span class="k"&gt;LIMIT&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Let's see how it works. We'll use SphinxQL and an explicit &lt;code&gt;GROUP BY&lt;/code&gt; - the new form of &lt;code&gt;GROUP_CONCAT()&lt;/code&gt; doesn't work without it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Events we'll work with
&lt;/h2&gt;

&lt;p&gt;Let's create an &lt;code&gt;activity&lt;/code&gt; table where each document is a single event. The event text is stored in &lt;code&gt;body&lt;/code&gt;, the user in &lt;code&gt;user_id&lt;/code&gt;, and the timestamp in &lt;code&gt;event_ts&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;activity&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;body&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;user_id&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;event_ts&lt;/span&gt; &lt;span class="nb"&gt;bigint&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;event_type&lt;/span&gt; &lt;span class="n"&gt;string&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;INSERT&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;activity&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;event_ts&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;event_type&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;VALUES&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1001&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'refund requested for order 501'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;       &lt;span class="mi"&gt;101&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1770000010&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'requested'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1002&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'user logged in'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;                       &lt;span class="mi"&gt;101&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1770000020&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'login'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1003&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'refund approved for order 501'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;        &lt;span class="mi"&gt;101&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1770000030&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'approved'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1004&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'refund email sent for order 501'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;      &lt;span class="mi"&gt;101&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1770000040&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'email'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1005&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'refund status checked for order 501'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="mi"&gt;101&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1770000050&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'checked'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1006&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'refund payout queued for order 501'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;   &lt;span class="mi"&gt;101&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1770000060&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'queued'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1007&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'refund webhook retried for order 501'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;101&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1770000060&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'retried'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2001&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'refund requested for order 601'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;       &lt;span class="mi"&gt;202&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1770000015&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'requested'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2002&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'refund approved for order 601'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;        &lt;span class="mi"&gt;202&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1770000025&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'approved'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2003&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'shipping address changed'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;             &lt;span class="mi"&gt;202&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1770000035&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'shipping'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2004&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'refund payout queued for order 601'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;   &lt;span class="mi"&gt;202&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1770000045&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'queued'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2005&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'refund completed for order 601'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;       &lt;span class="mi"&gt;202&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1770000055&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'completed'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;3001&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'refund requested for order 701'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;       &lt;span class="mi"&gt;303&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1770000012&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'requested'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;3002&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'invoice downloaded'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;                   &lt;span class="mi"&gt;303&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1770000022&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'invoice'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;3003&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'refund rejected for order 701'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;        &lt;span class="mi"&gt;303&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1770000032&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'rejected'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A search for &lt;code&gt;refund&lt;/code&gt; will find six events for user 101, four for user 202, and two for user 303. We deliberately gave events 1006 and 1007 the same timestamp: later you will see why sorting also needs &lt;code&gt;id&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's wrong with the old approaches
&lt;/h2&gt;

&lt;p&gt;Let's start with regular &lt;code&gt;GROUP_CONCAT()&lt;/code&gt;. The response format is what we need - one row per user:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt;
    &lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;matched_events&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;GROUP_CONCAT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;all_event_ids&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;activity&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="k"&gt;MATCH&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'refund'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;matched_events&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt; &lt;span class="k"&gt;ASC&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But for user 101, the string will contain all six IDs, for example &lt;code&gt;1001,1003,1004,1005,1006,1007&lt;/code&gt;, while we only need the five most recent ones. Also, without internal sorting, the order of the values is not guaranteed.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;+---------+----------------+-------------------------------+
| user_id | matched_events | all_event_ids                 |
+---------+----------------+-------------------------------+
|     101 |              6 | 1001,1003,1004,1005,1006,1007 |
|     202 |              4 | 2001,2002,2004,2005           |
|     303 |              2 | 3001,3003                     |
+---------+----------------+-------------------------------+
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Another option is to ask &lt;code&gt;GROUP N BY&lt;/code&gt; to select the five most recent documents from each group:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt;
    &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;event_ts&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;activity&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="k"&gt;MATCH&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'refund'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt;
&lt;span class="n"&gt;WITHIN&lt;/span&gt; &lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;event_ts&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt; &lt;span class="k"&gt;ASC&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The oldest event for user 101 disappears, but each remaining document is returned as a separate row. So instead of three rows, we get eleven: 5 + 4 + 2. This is convenient when the client needs the documents themselves, but it doesn't work for our compact summary.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;+------+---------+------------+
| id   | user_id | event_ts   |
+------+---------+------------+
| 1007 |     101 | 1770000060 |
| 1006 |     101 | 1770000060 |
| 1005 |     101 | 1770000050 |
| 1004 |     101 | 1770000040 |
| 1003 |     101 | 1770000030 |
| 2005 |     202 | 1770000055 |
| 2004 |     202 | 1770000045 |
| 2002 |     202 | 1770000025 |
| 2001 |     202 | 1770000015 |
| 3003 |     303 | 1770000032 |
| 3001 |     303 | 1770000012 |
+------+---------+------------+
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Combining only the five most recent IDs
&lt;/h2&gt;

&lt;p&gt;Now let's combine both steps: sort the documents directly inside &lt;code&gt;GROUP_CONCAT()&lt;/code&gt; and limit the list to five values there as well.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt;
    &lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;COUNT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;matched_events&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;GROUP_CONCAT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;id&lt;/span&gt;
        &lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;event_ts&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;
        &lt;span class="k"&gt;LIMIT&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;recent_event_ids&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;activity&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="k"&gt;MATCH&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'refund'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;matched_events&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt; &lt;span class="k"&gt;ASC&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;+---------+----------------+--------------------------+
| user_id | matched_events | recent_event_ids         |
+---------+----------------+--------------------------+
|     101 |              6 | 1007,1006,1005,1004,1003 |
|     202 |              4 | 2005,2004,2002,2001      |
|     303 |              2 | 3003,3001                |
+---------+----------------+--------------------------+
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;First, &lt;code&gt;MATCH('refund')&lt;/code&gt; selects refund events, then &lt;code&gt;GROUP BY user_id&lt;/code&gt; groups them by user. &lt;code&gt;COUNT(*)&lt;/code&gt; counts all matching events, while &lt;code&gt;GROUP_CONCAT()&lt;/code&gt; takes only the first five from each group after sorting.&lt;/p&gt;

&lt;p&gt;The second sort key - &lt;code&gt;id DESC&lt;/code&gt; - is especially important here. Events 1006 and 1007 have the same &lt;code&gt;event_ts&lt;/code&gt;, so without it their relative order would be undefined. Sorting by ID ensures that event 1007 always comes first.&lt;/p&gt;

&lt;p&gt;Notice that the query has two &lt;code&gt;ORDER BY&lt;/code&gt; clauses. The one inside &lt;code&gt;GROUP_CONCAT()&lt;/code&gt; determines the order of IDs in the string. The final &lt;code&gt;ORDER BY&lt;/code&gt; sorts the completed rows: first by the number of matches, then by &lt;code&gt;user_id&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The inner &lt;code&gt;LIMIT&lt;/code&gt; does not change &lt;code&gt;COUNT(*)&lt;/code&gt; or affect pagination of the overall result. So user 101 still has six matches even though only five IDs are shown next to it. That's exactly what we need for this summary.&lt;/p&gt;

&lt;p&gt;One more detail: &lt;code&gt;GROUP_CONCAT()&lt;/code&gt; always returns a string, even when it contains numeric IDs. If the API needs to return an array of numbers or objects with several fields, the string has to be parsed on the client side, or a different response format should be used.&lt;/p&gt;

&lt;p&gt;The query works the same way with distributed tables. Manticore collects candidates from all local and remote tables, then selects the overall top-N for each group.&lt;/p&gt;

&lt;h2&gt;
  
  
  If a comma doesn't work
&lt;/h2&gt;

&lt;p&gt;By default, values are separated by commas. Sometimes a more readable string is useful - for example, showing the event type next to its ID. You can set a custom separator with &lt;code&gt;SEPARATOR&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt;
    &lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;GROUP_CONCAT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;CONCAT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event_type&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;':'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;TO_STRING&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
        &lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;event_ts&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;
        &lt;span class="n"&gt;SEPARATOR&lt;/span&gt; &lt;span class="s1"&gt;' / '&lt;/span&gt;
        &lt;span class="k"&gt;LIMIT&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;recent_events&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;activity&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="k"&gt;MATCH&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'refund'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt; &lt;span class="k"&gt;ASC&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;+---------+----------------------------------------------+
| user_id | recent_events                                |
+---------+----------------------------------------------+
|     101 | retried:1007 / queued:1006 / checked:1005    |
|     202 | completed:2005 / queued:2004 / approved:2002 |
|     303 | rejected:3003 / requested:3001               |
+---------+----------------------------------------------+
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In this syntax, &lt;code&gt;SEPARATOR&lt;/code&gt; comes before &lt;code&gt;LIMIT&lt;/code&gt;. Manticore doesn't escape anything or add quotes: the result is a regular string, not a JSON array.&lt;/p&gt;

&lt;p&gt;This works well for &lt;code&gt;event_type&lt;/code&gt; because those values are controlled by the application. Be more careful with arbitrary text: if the separator appears in the data itself, the result can no longer be parsed reliably. In that case, it's better to return separate rows or use a structured format.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the new approach doesn't work
&lt;/h2&gt;

&lt;p&gt;This form of &lt;code&gt;GROUP_CONCAT()&lt;/code&gt; has several limitations. It works only in SQL queries with an explicit &lt;code&gt;GROUP BY&lt;/code&gt; and does not support:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;DISTINCT&lt;/code&gt;, &lt;code&gt;OFFSET&lt;/code&gt;, or combining multiple expressions at once;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;JOIN&lt;/code&gt;, &lt;code&gt;FACET&lt;/code&gt;, outer &lt;code&gt;SELECT&lt;/code&gt; queries, or table functions;&lt;/li&gt;
&lt;li&gt;KNN and hybrid queries, or scroll;&lt;/li&gt;
&lt;li&gt;implicit grouping or equivalent aggregation syntax in the JSON API.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Its alias cannot be used in &lt;code&gt;HAVING&lt;/code&gt; or the final &lt;code&gt;ORDER BY&lt;/code&gt;. The groups themselves can still be sorted by the grouping key and regular aggregates - for example, by &lt;code&gt;user_id&lt;/code&gt; and &lt;code&gt;COUNT(*)&lt;/code&gt;, as in the query above.&lt;/p&gt;

&lt;p&gt;Memory usage is another consideration. For each such expression, Manticore stores a separate top-N for every group that remains in the result. The more groups there are, the higher &lt;code&gt;N&lt;/code&gt; is, and the more expressions you use, the more memory is required. The size of the values and sort keys also affects memory usage, so it is best not to set an unnecessarily high limit "just in case".&lt;/p&gt;

&lt;p&gt;Full documentation for the new functionality &lt;a href="https://manual.manticoresearch.com/Searching/Grouping#GROUP_CONCAT%28field%29" rel="noopener noreferrer"&gt;is available here&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  More examples
&lt;/h2&gt;

&lt;p&gt;Support events are just one possible scenario. The new mode can be useful anywhere a document already contains a group key, the value you want to collect, and a field to sort by:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Product images:&lt;/strong&gt; collect the first few IDs or paths for each &lt;code&gt;product_id&lt;/code&gt;, sorted by &lt;code&gt;display_order&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Priority tasks:&lt;/strong&gt; return up to six task IDs for each &lt;code&gt;assignee_id&lt;/code&gt;, sorted by a precomputed &lt;code&gt;priority&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Server errors:&lt;/strong&gt; show the latest N errors for each server while keeping their total count separately.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you need a short list of IDs, names, or paths, &lt;code&gt;GROUP_CONCAT(... ORDER BY ... LIMIT N)&lt;/code&gt; now lets you get it in a single query. We hope you find this useful.&lt;/p&gt;

</description>
      <category>database</category>
      <category>sql</category>
      <category>opensource</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>How to search for long hashes and IDs with dict='keywords_32k'</title>
      <dc:creator>Sergey Nikolaev</dc:creator>
      <pubDate>Wed, 02 Sep 2026 07:30:01 +0000</pubDate>
      <link>https://dev.to/sanikolaev/how-to-search-for-long-hashes-and-ids-with-dictkeywords32k-oo2</link>
      <guid>https://dev.to/sanikolaev/how-to-search-for-long-hashes-and-ids-with-dictkeywords32k-oo2</guid>
      <description>&lt;p&gt;Originally posted by Manticore Search on &lt;a href="https://manticoresearch.com/blog/dict_keywords_32k/" rel="noopener noreferrer"&gt;https://manticoresearch.com/blog/dict_keywords_32k/&lt;/a&gt; on August 31, 2026&lt;/p&gt;

&lt;h1&gt;
  
  
  How to search for long hashes and IDs with dict='keywords_32k'
&lt;/h1&gt;

&lt;p&gt;A practical guide to searching long hashes, event IDs, message IDs, and email addresses in Manticore Search: limits, exact matching, wildcard search, tokenization, migration, and limitations.&lt;/p&gt;

&lt;p&gt;Full-text search usually works with ordinary words: product names, titles, comments, and descriptions. Such tokens are rarely longer than a few dozen characters.&lt;/p&gt;

&lt;p&gt;Logs and technical data are different. A SHA-256 hash is 64 characters long, while message IDs, correlation keys, event identifiers, and some email addresses can be even longer. The value is often meaningful only as a whole: if its tail is lost, one ID can easily be confused with another.&lt;/p&gt;

&lt;p&gt;Manticore Search provides &lt;code&gt;dict='keywords_32k'&lt;/code&gt; for these cases.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;keywords_32k&lt;/code&gt; is available starting with Manticore Search 27.1.1. Version 27.1.5 or newer is recommended when converting an existing table from &lt;code&gt;keywords&lt;/code&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The problem with the regular dictionary
&lt;/h2&gt;

&lt;p&gt;By default, Manticore uses &lt;code&gt;dict='keywords'&lt;/code&gt;. Its maximum token length is &lt;strong&gt;42 bytes after normalization&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The limit is measured in bytes, not characters. One ASCII character takes one byte, but a UTF-8 character can take several bytes.&lt;/p&gt;

&lt;p&gt;If a token exceeds 42 bytes, Manticore truncates it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;when indexing a document;&lt;/li&gt;
&lt;li&gt;when processing a search query.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;As a result, querying the complete long value does not necessarily return zero results. Because the query is truncated too, the document may be found-but only by the first 42 bytes.&lt;/p&gt;

&lt;p&gt;This creates a more serious problem: two different IDs with the same first 42 bytes become indistinguishable to full-text search. It is also impossible to find a value by a fragment located after the 42nd byte.&lt;/p&gt;

&lt;h2&gt;
  
  
  What &lt;code&gt;keywords_32k&lt;/code&gt; changes
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;dict='keywords_32k'&lt;/code&gt; increases the maximum normalized token length to &lt;strong&gt;32768 bytes&lt;/strong&gt;, or 32 KB.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Behavior&lt;/th&gt;
&lt;th&gt;&lt;code&gt;dict='keywords'&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;&lt;code&gt;dict='keywords_32k'&lt;/code&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Maximum token length&lt;/td&gt;
&lt;td&gt;42 bytes&lt;/td&gt;
&lt;td&gt;32768 bytes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Token exceeding the limit&lt;/td&gt;
&lt;td&gt;Truncated&lt;/td&gt;
&lt;td&gt;Skipped with a warning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prefix and infix search&lt;/td&gt;
&lt;td&gt;Supported&lt;/td&gt;
&lt;td&gt;Supported&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Morphology for tokens longer than 42 bytes&lt;/td&gt;
&lt;td&gt;Token is already truncated&lt;/td&gt;
&lt;td&gt;Not applied&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RT tables&lt;/td&gt;
&lt;td&gt;Supported&lt;/td&gt;
&lt;td&gt;Supported&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Plain tables&lt;/td&gt;
&lt;td&gt;Supported&lt;/td&gt;
&lt;td&gt;Supported&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The setting applies to the entire table:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;events&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;message&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;event_id&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;dict&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'keywords_32k'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Regular short words in the same table continue to use the configured morphology. Tokens longer than 42 bytes are stored in their original normalized form, without stemming or lemmatization.&lt;/p&gt;

&lt;p&gt;For machine identifiers, this is usually exactly what you need: a hash or event ID has no useful word stem.&lt;/p&gt;

&lt;h2&gt;
  
  
  When to use &lt;code&gt;keywords_32k&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;Use it when both conditions are true:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The value can exceed 42 bytes after tokenization.&lt;/li&gt;
&lt;li&gt;It must be searchable through &lt;code&gt;MATCH()&lt;/code&gt;, by prefix, or by substring.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Typical examples include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;SHA-256 and other long hashes;&lt;/li&gt;
&lt;li&gt;event IDs and message IDs;&lt;/li&gt;
&lt;li&gt;request IDs, trace IDs, and other technical identifiers;&lt;/li&gt;
&lt;li&gt;long record keys;&lt;/li&gt;
&lt;li&gt;email addresses with a long local or domain part;&lt;/li&gt;
&lt;li&gt;technical values from logs;&lt;/li&gt;
&lt;li&gt;long identifiers containing separators.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;However, &lt;code&gt;keywords_32k&lt;/code&gt; is not required for every ID lookup.&lt;/p&gt;

&lt;h3&gt;
  
  
  If you only need exact equality
&lt;/h3&gt;

&lt;p&gt;If the application always receives the complete ID and only needs to check exact equality, a string attribute is sufficient:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;events&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;message&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;event_id&lt;/span&gt; &lt;span class="n"&gt;string&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then use a regular filter:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;message&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;events&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;event_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
  &lt;span class="s1"&gt;'9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  If you need both exact matching and full-text search
&lt;/h3&gt;

&lt;p&gt;Use &lt;code&gt;string attribute indexed&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;events&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;message&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;event_id&lt;/span&gt; &lt;span class="n"&gt;string&lt;/span&gt; &lt;span class="n"&gt;attribute&lt;/span&gt; &lt;span class="n"&gt;indexed&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;dict&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'keywords_32k'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In this case, Manticore:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;stores the original value as a string attribute;&lt;/li&gt;
&lt;li&gt;lets you filter it with &lt;code&gt;WHERE&lt;/code&gt;;&lt;/li&gt;
&lt;li&gt;also indexes it for &lt;code&gt;MATCH()&lt;/code&gt; and wildcard search.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is usually the most convenient schema for technical identifiers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Searching by the complete token
&lt;/h2&gt;

&lt;p&gt;Create a table and add a 64-character SHA-256 hash:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;DROP&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;IF&lt;/span&gt; &lt;span class="k"&gt;EXISTS&lt;/span&gt; &lt;span class="n"&gt;events&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;events&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;message&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;event_id&lt;/span&gt; &lt;span class="n"&gt;string&lt;/span&gt; &lt;span class="n"&gt;attribute&lt;/span&gt; &lt;span class="n"&gt;indexed&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;dict&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'keywords_32k'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;INSERT&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;events&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;event_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;VALUES&lt;/span&gt;
&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="s1"&gt;'delivery accepted'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="s1"&gt;'9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08'&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use the complete normalized token in a full-text search:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;message&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;events&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="k"&gt;MATCH&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="s1"&gt;'@event_id 9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08'&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;@event_id&lt;/code&gt; operator restricts the search to the required field. Without it, Manticore searches for the value in all full-text fields of the table.&lt;/p&gt;

&lt;p&gt;Quotes are not needed around a single simple alphanumeric token here. Quotes denote a phrase search; they do not turn &lt;code&gt;MATCH()&lt;/code&gt; into a byte-for-byte comparison of the original string.&lt;/p&gt;

&lt;h2&gt;
  
  
  A full-text match is not the same as exact equality
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;MATCH()&lt;/code&gt; operates on the result of tokenization and normalization. It can be affected by:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;charset_table&lt;/code&gt;;&lt;/li&gt;
&lt;li&gt;conversion to lowercase;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;blend_chars&lt;/code&gt;;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;ignore_chars&lt;/code&gt;;&lt;/li&gt;
&lt;li&gt;word forms and other text-processing settings.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For a strict comparison of the stored value, use a string attribute:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;collation_connection&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'binary'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;message&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;events&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;event_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
  &lt;span class="s1"&gt;'9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;binary&lt;/code&gt; enables byte-for-byte string comparisons in the current SQL session. This setting does not affect full-text search behavior.&lt;/p&gt;

&lt;p&gt;The practical rule is simple:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;WHERE event_id = ...&lt;/code&gt; - a strict comparison of the stored string;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;MATCH('@event_id ...')&lt;/code&gt; - a normalized-token search;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;MATCH('@event_id prefix*')&lt;/code&gt; - a prefix search;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;MATCH('@event_id *fragment*')&lt;/code&gt; - a substring search.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How to inspect tokenization
&lt;/h2&gt;

&lt;p&gt;Before loading a large volume of data, check that Manticore actually sees the value as a single token:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CALL&lt;/span&gt; &lt;span class="n"&gt;KEYWORDS&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="s1"&gt;'9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="s1"&gt;'events'&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;normalized&lt;/code&gt; column should contain the complete 64-character hash.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;CALL KEYWORDS&lt;/code&gt; is especially useful for values containing periods, hyphens, &lt;code&gt;@&lt;/code&gt;, colons, slashes, and characters from different writing systems. This lets you inspect the actual token boundaries before indexing the data.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prefix and substring search
&lt;/h2&gt;

&lt;p&gt;To search inside a token, enable &lt;code&gt;min_infix_len&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;DROP&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;IF&lt;/span&gt; &lt;span class="k"&gt;EXISTS&lt;/span&gt; &lt;span class="n"&gt;events_infix&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;events_infix&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;message&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;event_id&lt;/span&gt; &lt;span class="n"&gt;string&lt;/span&gt; &lt;span class="n"&gt;attribute&lt;/span&gt; &lt;span class="n"&gt;indexed&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;dict&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'keywords_32k'&lt;/span&gt;
&lt;span class="n"&gt;min_infix_len&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'4'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Prefix search:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;message&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;events_infix&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="k"&gt;MATCH&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'@event_id 9f86d081*'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Substring search:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;message&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;events_infix&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="k"&gt;MATCH&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'@event_id *b2b0b822*'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A positive &lt;code&gt;min_infix_len&lt;/code&gt; also enables prefix search. This example uses &lt;code&gt;4&lt;/code&gt; to prevent excessively short and broad patterns.&lt;/p&gt;

&lt;p&gt;For production systems: do not allow excessively short fragments, choose &lt;code&gt;min_infix_len&lt;/code&gt; based on real data, use &lt;code&gt;expansion_limit&lt;/code&gt;, test performance with a production-sized dictionary, and restrict search to a specific field with &lt;code&gt;@field&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Email addresses and other values with separators
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;keywords_32k&lt;/code&gt; changes only the maximum token length. It does not determine where a token begins and ends.&lt;/p&gt;

&lt;p&gt;By default, a period, &lt;code&gt;@&lt;/code&gt;, hyphen, and other characters may split a value into multiple parts. If an email address or message ID must also be indexed as a whole, you can use &lt;code&gt;blend_chars&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;DROP&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;IF&lt;/span&gt; &lt;span class="k"&gt;EXISTS&lt;/span&gt; &lt;span class="n"&gt;mail_events&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;mail_events&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;sender&lt;/span&gt; &lt;span class="n"&gt;string&lt;/span&gt; &lt;span class="n"&gt;attribute&lt;/span&gt; &lt;span class="n"&gt;indexed&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;subject&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;dict&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'keywords_32k'&lt;/span&gt;
&lt;span class="n"&gt;blend_chars&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'., @, -'&lt;/span&gt;
&lt;span class="n"&gt;min_infix_len&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'4'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Blended characters are indexed in two ways: as part of the complete token and as separators between regular parts of the value. This makes it possible to search both the complete email address and individual words within it.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to convert an existing table
&lt;/h2&gt;

&lt;p&gt;For an RT table, you can change the setting with &lt;code&gt;ALTER TABLE&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;events&lt;/span&gt; &lt;span class="n"&gt;dict&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'keywords_32k'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;However, this affects only documents added or replaced after the setting is changed. Existing documents are not automatically tokenized again. Their long tokens remain in the old truncated form until the documents are reindexed.&lt;/p&gt;

&lt;p&gt;Follow these steps:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Update &lt;code&gt;dict&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Verify the setting with &lt;code&gt;SHOW CREATE TABLE&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Reindex or reload the existing documents.&lt;/li&gt;
&lt;li&gt;Check several long values with &lt;code&gt;CALL KEYWORDS&lt;/code&gt; and &lt;code&gt;MATCH()&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Why &lt;code&gt;dict='crc'&lt;/code&gt; does not solve this problem
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;dict='crc'&lt;/code&gt; stores keyword checksums instead of their original text. However, it does not increase the allowed token length. The exception to the regular 42-byte limit is implemented specifically by &lt;code&gt;dict='keywords_32k'&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Current limitations
&lt;/h2&gt;

&lt;p&gt;At the time of publication, &lt;code&gt;dict='keywords_32k'&lt;/code&gt; has several limitations:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;CALL SUGGEST&lt;/code&gt; and &lt;code&gt;CALL QSUGGEST&lt;/code&gt; are not supported;&lt;/li&gt;
&lt;li&gt;it cannot be used in percolate tables;&lt;/li&gt;
&lt;li&gt;tokens longer than 42 bytes are not highlighted in snippets or highlights;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;indextool --dumpdict&lt;/code&gt; cannot dump this type of dictionary;&lt;/li&gt;
&lt;li&gt;the full-text &lt;code&gt;REGEX&lt;/code&gt; operator works with &lt;code&gt;dict='keywords'&lt;/code&gt;, but not with &lt;code&gt;keywords_32k&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Do not index secrets
&lt;/h2&gt;

&lt;p&gt;Being able to search a long value does not mean that you should store it in a search index. Do not index API keys, bearer tokens, session cookies, private keys, passwords, reset tokens, or other data that grants access to the system unless absolutely necessary.&lt;/p&gt;

&lt;p&gt;If a secret must be matched by its exact value, it is safer to calculate a suitable hash in advance and store only the hash.&lt;/p&gt;

&lt;h2&gt;
  
  
  Quick checklist
&lt;/h2&gt;

&lt;p&gt;Before enabling &lt;code&gt;keywords_32k&lt;/code&gt;, check:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Is the normalized token actually longer than 42 bytes?&lt;/li&gt;
&lt;li&gt;Do you need full-text or wildcard search rather than only &lt;code&gt;WHERE value = ...&lt;/code&gt;?&lt;/li&gt;
&lt;li&gt;Does &lt;code&gt;CALL KEYWORDS&lt;/code&gt; see the entire value as one token?&lt;/li&gt;
&lt;li&gt;Do you need &lt;code&gt;blend_chars&lt;/code&gt; for periods, hyphens, &lt;code&gt;@&lt;/code&gt;, and other separators?&lt;/li&gt;
&lt;li&gt;Is &lt;code&gt;min_infix_len&lt;/code&gt; large enough?&lt;/li&gt;
&lt;li&gt;Is the number of wildcard expansions limited?&lt;/li&gt;
&lt;li&gt;Have the old documents been reindexed?&lt;/li&gt;
&lt;li&gt;Is the field free of secret data?&lt;/li&gt;
&lt;li&gt;Does the application avoid relying on highlighting, &lt;code&gt;SUGGEST&lt;/code&gt;, percolate, or full-text &lt;code&gt;REGEX&lt;/code&gt;?&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Summary
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;dict='keywords_32k'&lt;/code&gt; solves a specific problem: it allows a full-text index to store normalized tokens up to 32768 bytes long instead of the regular 42 bytes.&lt;/p&gt;

&lt;p&gt;It is well suited to long hashes, event IDs, message IDs, email addresses, and other machine identifiers. Keep three points in mind:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;keywords_32k&lt;/code&gt; increases the token length but does not change tokenization rules;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;MATCH()&lt;/code&gt; on a complete token is not the same as a strict comparison of the original string;&lt;/li&gt;
&lt;li&gt;existing documents must be reindexed after the setting is changed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you only need exact equality, use a string attribute. If you need both exact matching and partial-value search, use &lt;code&gt;string attribute indexed&lt;/code&gt; together with &lt;code&gt;dict='keywords_32k'&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Documentation
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://manual.manticoresearch.com/Creating_a_table/NLP_and_tokenization/Low-level_tokenization#dict" rel="noopener noreferrer"&gt;&lt;code&gt;dict&lt;/code&gt; and &lt;code&gt;keywords_32k&lt;/code&gt; limitations&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://manual.manticoresearch.com/Creating_a_table/NLP_and_tokenization/Data_tokenization#Token-length-limit" rel="noopener noreferrer"&gt;Token length limit&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://manual.manticoresearch.com/Creating_a_table/NLP_and_tokenization/Wildcard_searching_settings" rel="noopener noreferrer"&gt;Wildcard search settings&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://manual.manticoresearch.com/Creating_a_table/NLP_and_tokenization/Low-level_tokenization#blend_chars" rel="noopener noreferrer"&gt;&lt;code&gt;blend_chars&lt;/code&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://manual.manticoresearch.com/Searching/Autocomplete#CALL-KEYWORDS" rel="noopener noreferrer"&gt;&lt;code&gt;CALL KEYWORDS&lt;/code&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://manual.manticoresearch.com/Creating_a_table/Local_tables#String" rel="noopener noreferrer"&gt;String attributes and indexed strings&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://manual.manticoresearch.com/Updating_table_schema_and_settings" rel="noopener noreferrer"&gt;Updating full-text settings and reindexing&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://manual.manticoresearch.com/Searching/Collations" rel="noopener noreferrer"&gt;Collations and string comparison&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://manticoresearch.com/blog/manticore-search-27-1-5/" rel="noopener noreferrer"&gt;Manticore Search 27.1.5&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>database</category>
      <category>search</category>
      <category>sql</category>
      <category>performance</category>
    </item>
    <item>
      <title>How KupujemProdajem searches 5.6 million ads at ~10 ms with Manticore Search</title>
      <dc:creator>Sergey Nikolaev</dc:creator>
      <pubDate>Mon, 31 Aug 2026 07:00:01 +0000</pubDate>
      <link>https://dev.to/sanikolaev/how-kupujemprodajem-searches-56-million-ads-at-10-ms-with-manticore-search-1373</link>
      <guid>https://dev.to/sanikolaev/how-kupujemprodajem-searches-56-million-ads-at-10-ms-with-manticore-search-1373</guid>
      <description>&lt;p&gt;Originally posted by Manticore Search on &lt;a href="https://manticoresearch.com/blog/kupujemprodajem/" rel="noopener noreferrer"&gt;https://manticoresearch.com/blog/kupujemprodajem/&lt;/a&gt; on August 31, 2026&lt;/p&gt;

&lt;h1&gt;
  
  
  How KupujemProdajem searches 5.6 million ads at ~10 ms with Manticore Search
&lt;/h1&gt;

&lt;p&gt;How Serbia's largest online classifieds marketplace uses Manticore Search to search 5.6 million active ads at ~10 ms, process real-time updates, and prepares for hybrid search.&lt;/p&gt;

&lt;p&gt;For an online marketplace, having millions of listings is useful only if buyers can find the right one.&lt;/p&gt;

&lt;p&gt;At &lt;a href="https://www.kupujemprodajem.com/" rel="noopener noreferrer"&gt;KupujemProdajem&lt;/a&gt;, Serbia's largest online classifieds marketplace, search is the primary way users discover ads. Buyers search from the web and mobile applications, while internal teams use search for moderation and support. Because of that, both relevance and latency directly affect user engagement and sellers' ability to be discovered.&lt;/p&gt;

&lt;p&gt;Today, KupujemProdajem uses Manticore Search to search approximately &lt;strong&gt;5.6 million active ads&lt;/strong&gt;, while separate administrative indexes contain around &lt;strong&gt;70 million rows&lt;/strong&gt; of current and historical data. The production cluster handles around &lt;strong&gt;150 search requests per second&lt;/strong&gt; on average, with an average Manticore query time of approximately &lt;strong&gt;10 ms&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;And the same search infrastructure is now giving the team a path toward something more ambitious: &lt;a href="https://manticoresearch.com/blog/hybrid-search/" rel="noopener noreferrer"&gt;hybrid lexical and semantic search&lt;/a&gt; designed specifically for Serbian-language marketplace content.&lt;/p&gt;

&lt;h2&gt;
  
  
  Search is part of the core product
&lt;/h2&gt;

&lt;p&gt;KupujemProdajem maintains two main kinds of search indexes. The first powers the customer-facing experience. It contains approximately 5.6 million active classified ads, including titles, descriptions, categories, and structured attributes. Users search these indexes while browsing the marketplace on web and mobile.&lt;/p&gt;

&lt;p&gt;The second set of indexes is used internally. In addition to current data, these indexes retain part of the history of deleted ads, allowing moderation and support teams to search information that may no longer be visible on the public marketplace.&lt;/p&gt;

&lt;p&gt;This makes search important on both sides of the product. For buyers, it determines how quickly they can get from an idea of what they want to a relevant listing. For sellers, search affects whether their ads are discovered. Internally, it gives employees access to historical marketplace data when investigating problems.&lt;/p&gt;

&lt;p&gt;As &lt;a href="https://www.linkedin.com/in/dimitrije-marinkovi%C4%87-131bb62bb/" rel="noopener noreferrer"&gt;Dimitrije Marinković&lt;/a&gt;, Backend Developer at KupujemProdajem, puts it:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Search is the primary discovery mechanism on the platform, so relevance and latency directly affect user engagement and seller success.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Choosing a search engine that could grow with the product
&lt;/h2&gt;

&lt;p&gt;KupujemProdajem had already been running a dedicated search engine for years. When its previous engine, Sphinx, stopped being actively developed, the team needed a maintained alternative that would fit the existing architecture without requiring the entire search integration to be rebuilt.&lt;/p&gt;

&lt;p&gt;But compatibility was only part of the decision. The new engine also needed to support:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;consistently low query latency;&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://manual.manticoresearch.com/Creating_a_table/Local_tables/Real-time_table" rel="noopener noreferrer"&gt;real-time changes&lt;/a&gt; to millions of ads;&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://manual.manticoresearch.com/Creating_a_cluster/Setting_up_replication/Setting_up_replication" rel="noopener noreferrer"&gt;replication&lt;/a&gt;;&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://manual.manticoresearch.com/Connecting_to_the_server/MySQL_protocol" rel="noopener noreferrer"&gt;SQL&lt;/a&gt; that worked naturally with the team's MySQL-centric PHP stack;&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://manual.manticoresearch.com/Searching/Full_text_matching/Basic_usage" rel="noopener noreferrer"&gt;full-text search&lt;/a&gt; together with &lt;a href="https://manual.manticoresearch.com/Searching/Filters" rel="noopener noreferrer"&gt;structured filters&lt;/a&gt;;&lt;/li&gt;
&lt;li&gt;flexible &lt;a href="https://manual.manticoresearch.com/Searching/Options#ranker" rel="noopener noreferrer"&gt;relevance tuning&lt;/a&gt;;&lt;/li&gt;
&lt;li&gt;and a path toward &lt;a href="https://manticoresearch.com/blog/vector-search-deep-dive/" rel="noopener noreferrer"&gt;vector&lt;/a&gt; and &lt;a href="https://manual.manticoresearch.com/Searching/Hybrid_search" rel="noopener noreferrer"&gt;hybrid search&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Manticore provided those capabilities in one system. That last requirement is becoming increasingly important.&lt;/p&gt;

&lt;p&gt;KupujemProdajem plans to combine traditional lexical search with vector search using its own &lt;a href="https://manticoresearch.com/blog/vector-search-in-databases/#embeddings" rel="noopener noreferrer"&gt;embedding model&lt;/a&gt;. Because Manticore supports &lt;a href="https://manual.manticoresearch.com/Searching/KNN" rel="noopener noreferrer"&gt;KNN search&lt;/a&gt; alongside full-text search, the team can build that functionality inside the search infrastructure it already operates instead of synchronizing its data with a separate vector database.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keeping millions of changing ads searchable in real time
&lt;/h2&gt;

&lt;p&gt;MySQL remains the source of truth for marketplace data. When an ad is created, edited, or removed, the corresponding change is pushed to Manticore in real time. The production workload averages approximately &lt;strong&gt;40 &lt;a href="https://manual.manticoresearch.com/Data_creation_and_modification/Updating_documents/REPLACE" rel="noopener noreferrer"&gt;&lt;code&gt;REPLACE&lt;/code&gt;&lt;/a&gt; operations per second&lt;/strong&gt;, with additional &lt;a href="https://manual.manticoresearch.com/Data_creation_and_modification/Updating_documents/UPDATE" rel="noopener noreferrer"&gt;&lt;code&gt;UPDATE&lt;/code&gt;&lt;/a&gt; and &lt;a href="https://manual.manticoresearch.com/Data_creation_and_modification/Deleting_documents" rel="noopener noreferrer"&gt;&lt;code&gt;DELETE&lt;/code&gt;&lt;/a&gt; operations.&lt;/p&gt;

&lt;p&gt;This is particularly useful for classifieds. Marketplace inventory changes continuously. Sellers publish new ads, update prices and descriptions, change status, and remove products that are no longer available.&lt;/p&gt;

&lt;p&gt;A search system therefore has two jobs at once:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;return relevant results quickly;&lt;/li&gt;
&lt;li&gt;make sure those results reflect the current marketplace.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;KupujemProdajem uses Manticore's real-time updates to keep the searchable representation of the marketplace synchronized with MySQL as these changes happen.&lt;/p&gt;

&lt;h2&gt;
  
  
  More than matching words
&lt;/h2&gt;

&lt;p&gt;Marketplace search also has very different requirements from searching a collection of plain documents. A listing contains text, but it also has structured information such as category, price, location, status, and other properties. KupujemProdajem combines full-text matching with filters over these typed attributes in Manticore.&lt;/p&gt;

&lt;p&gt;The team also controls how different parts of the listing affect relevance. The title, for example, receives significantly more weight than the description or category. This makes intuitive sense for classified ads: a product named directly in the title is normally a stronger signal than the same word appearing somewhere in a longer description.&lt;/p&gt;

&lt;p&gt;On top of this, KupujemProdajem uses a &lt;a href="https://manual.manticoresearch.com/Searching/Sorting_and_ranking#Ranking-overview" rel="noopener noreferrer"&gt;custom ranker&lt;/a&gt; tuned for its relevance requirements and &lt;a href="https://manual.manticoresearch.com/Searching/Grouping" rel="noopener noreferrer"&gt;groups results&lt;/a&gt; by ad. This gives the team the ability to encode marketplace-specific relevance rules rather than relying entirely on a generic ranking formula.&lt;/p&gt;

&lt;h2&gt;
  
  
  A small replicated cluster
&lt;/h2&gt;

&lt;p&gt;The entire Manticore deployment runs on four virtual machines. Three nodes form the production replication cluster, while a fourth VM serves as a backup. All writes are sent to one node and replicated to the others. Read requests are distributed among all three production nodes by application-level logic.&lt;/p&gt;

&lt;p&gt;The architecture is relatively simple: &lt;strong&gt;MySQL → real-time updates → Manticore write node → replicated Manticore nodes → web/mobile/admin search&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That simplicity matters operationally. The team does not need a large distributed data platform merely to support marketplace search. The same cluster handles ingestion, full-text retrieval, structured filtering, custom ranking, replication, and the foundation for future vector search.&lt;/p&gt;

&lt;h2&gt;
  
  
  5.6 million active ads, 70 million admin rows
&lt;/h2&gt;

&lt;p&gt;The user-facing indexes currently contain around &lt;strong&gt;5.6 million active ads&lt;/strong&gt; and occupy approximately &lt;strong&gt;9 GB&lt;/strong&gt;. The larger administrative indexes contain approximately &lt;strong&gt;70 million rows&lt;/strong&gt;, including some historical deleted-ad data, and occupy around &lt;strong&gt;90 GB&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Across the three-node cluster, KupujemProdajem sees approximately:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Production scale&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Active ads&lt;/td&gt;
&lt;td&gt;~5.6 million&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Admin index rows&lt;/td&gt;
&lt;td&gt;~70 million&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;User-facing index size&lt;/td&gt;
&lt;td&gt;~9 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Admin index size&lt;/td&gt;
&lt;td&gt;~90 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Search traffic&lt;/td&gt;
&lt;td&gt;~150 queries/sec average&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Search traffic per node&lt;/td&gt;
&lt;td&gt;~50 queries/sec average&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;REPLACE&lt;/code&gt; traffic&lt;/td&gt;
&lt;td&gt;~40/sec&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Average query time&lt;/td&gt;
&lt;td&gt;~10 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Infrastructure&lt;/td&gt;
&lt;td&gt;3 production VMs + 1 backup&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The cluster remains stable under normal operation, without stuck or long-running queries. Short latency spikes of approximately 150 ms occur occasionally, while the average query execution time remains around 10 ms. Those numbers are particularly useful because this is not a synthetic benchmark. They describe an actual production workload behind the primary discovery experience of a large marketplace.&lt;/p&gt;

&lt;h2&gt;
  
  
  What users get from this architecture
&lt;/h2&gt;

&lt;p&gt;Infrastructure metrics matter, but only because of what they enable at the product level. For buyers, Manticore helps KupujemProdajem keep search fast even while querying millions of listings and applying relevance logic and structured filters. Because changes are continuously pushed into the index, search can also reflect changes in marketplace inventory quickly instead of relying on occasional large reindexing jobs.&lt;/p&gt;

&lt;p&gt;For sellers, this infrastructure supports the main mechanism through which buyers discover their listings. For moderation and support teams, the much larger internal indexes make historical ad information searchable even when it is no longer part of the live marketplace.&lt;/p&gt;

&lt;h2&gt;
  
  
  The next challenge: understanding Serbian-language intent
&lt;/h2&gt;

&lt;p&gt;The next project for the KupujemProdajem team is &lt;a href="https://manticoresearch.com/blog/hybrid-search/" rel="noopener noreferrer"&gt;hybrid search&lt;/a&gt;. Traditional full-text search works particularly well when the words entered by a user also appear in an ad. But marketplace users and sellers do not always describe the same thing in exactly the same way.&lt;/p&gt;

&lt;p&gt;That problem becomes even more interesting in Serbian. The team highlights several challenges:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;rich &lt;a href="https://manual.manticoresearch.com/Creating_a_table/NLP_and_tokenization/Morphology" rel="noopener noreferrer"&gt;morphology&lt;/a&gt;;&lt;/li&gt;
&lt;li&gt;both Latin and Cyrillic scripts;&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://manual.manticoresearch.com/Creating_a_table/NLP_and_tokenization/Wordforms" rel="noopener noreferrer"&gt;synonyms&lt;/a&gt;;&lt;/li&gt;
&lt;li&gt;different ways buyers and sellers can express the same intent.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;KupujemProdajem is therefore preparing its own embedding model for Serbian marketplace content. The plan is to retrieve results using both &lt;a href="https://manticoresearch.com/use-case/lexical-search/" rel="noopener noreferrer"&gt;lexical matching&lt;/a&gt; and KNN vector search, then combine the two sets of signals at query time.&lt;/p&gt;

&lt;p&gt;The user benefit is straightforward. Someone may search for an item using words that never literally appear in the seller's listing, while the listing is nevertheless highly relevant. Semantic retrieval can provide another signal for recognizing that relationship.&lt;/p&gt;

&lt;p&gt;As Dimitrije describes the expected result:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Semantic matching will let users find relevant ads even when their wording doesn't literally match the listing.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Perhaps just as important for the engineering team, they do not need to introduce another database to do it. The vector representation can live alongside the text and structured attributes already searchable in Manticore, giving the team one engine in which to combine lexical relevance, filters, and semantic similarity.&lt;/p&gt;

&lt;h2&gt;
  
  
  A foundation for better marketplace search
&lt;/h2&gt;

&lt;p&gt;KupujemProdajem already uses Manticore for one of the most important workflows on its marketplace. Approximately &lt;strong&gt;5.6 million active ads&lt;/strong&gt; are searchable with an average engine query time of &lt;strong&gt;around 10 ms&lt;/strong&gt;. Updates arrive continuously. Replication provides multiple search nodes. Internal teams search a much larger historical dataset using the same technology.&lt;/p&gt;

&lt;p&gt;But the more interesting part may be what comes next. The team can now experiment with semantic retrieval and hybrid relevance without replacing the search infrastructure or operating an additional vector database.&lt;/p&gt;

&lt;p&gt;For KupujemProdajem, that means Manticore is not only keeping today's marketplace search fast. It is also giving the team a practical way to make tomorrow's search better.&lt;/p&gt;

</description>
      <category>database</category>
      <category>search</category>
      <category>performance</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Full-Text Search Still Works. It Just Doesn’t Get You to an Answer</title>
      <dc:creator>Sergey Nikolaev</dc:creator>
      <pubDate>Mon, 24 Aug 2026 10:34:39 +0000</pubDate>
      <link>https://dev.to/sanikolaev/full-text-search-still-works-it-just-doesnt-get-you-to-an-answer-1dhd</link>
      <guid>https://dev.to/sanikolaev/full-text-search-still-works-it-just-doesnt-get-you-to-an-answer-1dhd</guid>
      <description>&lt;p&gt;Originally posted by Klim Todrik on &lt;a href="https://manticoresearch.com/blog/conversational-search/" rel="noopener noreferrer"&gt;https://manticoresearch.com/blog/conversational-search/&lt;/a&gt; on Aug 20, 2026&lt;/p&gt;

&lt;p&gt;Imagine a typical online shoe store.&lt;/p&gt;

&lt;p&gt;A shopper opens search and types:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;I need black waterproof running shoes for daily runs on wet pavement. What would you recommend?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A few years ago, almost no one expected this from a search box. The query would have been shortened to something like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;black waterproof running shoes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then the shopper would open several product cards, compare descriptions, materials, intended use, and price, and make the decision alone.&lt;/p&gt;

&lt;p&gt;Today, people increasingly expect search itself to do part of that work. Not because full-text search has become worse. It still performs very well with exact names, SKUs, product codes, brands, and keywords. What has changed is what people expect to be able to ask a search system.&lt;/p&gt;

&lt;p&gt;For example, Google reported in May 2026 that AI Mode had surpassed one billion monthly users. The company notes that people are asking longer, more complex questions that previously did not fit into conventional search. (&lt;a href="https://blog.google/products-and-platforms/products/search/search-io-2026/" rel="noopener noreferrer"&gt;blog.google&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;The same thing happens in an online store.&lt;/p&gt;

&lt;p&gt;A query such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;I need black waterproof running shoes for daily runs on wet pavement.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;only looks like one sentence. For a search system, it contains several tasks.&lt;/p&gt;

&lt;p&gt;It needs to extract constraints such as color and waterproofing; understand that this is about running rather than walking; account for price, size, and availability; find suitable products; and, if the user asks “which are better,” explain the differences.&lt;/p&gt;

&lt;p&gt;No single algorithm solves all of that.&lt;/p&gt;

&lt;p&gt;Full-text search, vector search, filters, hybrid search, and a language model each solve different parts of the problem. It is far more effective to combine them than to choose between them.&lt;/p&gt;

&lt;p&gt;That is exactly why Manticore has Conversational Search.&lt;/p&gt;




&lt;h2&gt;
  
  
  Word search, semantic search, and conversation are different tasks
&lt;/h2&gt;

&lt;p&gt;Start with a simple query:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Nike Pegasus 41 black
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here, the system barely needs to interpret the user’s intent. Full-text search handles it directly.&lt;/p&gt;

&lt;p&gt;Or something even simpler:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SKU 123456
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Semantic methods are not needed here.&lt;/p&gt;

&lt;p&gt;Now consider another example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;light shoes for long summer walks
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A product card may not contain the words “summer” or “long walks,” but it may include details such as “breathable material” or “lightweight construction.”&lt;/p&gt;

&lt;p&gt;This is where vector search becomes useful.&lt;/p&gt;

&lt;p&gt;Real queries often fall between these extremes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;black Gore-Tex shoes for everyday running
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Some parameters — &lt;code&gt;black&lt;/code&gt; and &lt;code&gt;Gore-Tex&lt;/code&gt; — need to be preserved. &lt;code&gt;Everyday running&lt;/code&gt; describes the user’s intent rather than an exact attribute.&lt;/p&gt;

&lt;p&gt;For such cases, Manticore uses hybrid search, combining full-text and vector search through result ranking.&lt;/p&gt;

&lt;p&gt;But even hybrid search returns only a list of results.&lt;/p&gt;

&lt;p&gt;At that point, search considers its job done. The user usually does not.&lt;/p&gt;

&lt;p&gt;It does not answer questions such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Which of these models are better suited to rain?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And it certainly does not handle a follow-up such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Which of those cost less than $120?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is a conversation. It needs another layer.&lt;/p&gt;




&lt;h2&gt;
  
  
  What we built
&lt;/h2&gt;

&lt;p&gt;To test this in practice, we used &lt;a href="https://arxiv.org/abs/2602.16938" rel="noopener noreferrer"&gt;ConvApparel&lt;/a&gt;, a dataset of conversations about choosing apparel. After cleanup, it contained 82,524 products: footwear, pants, tops, and outerwear. Each product has a description, category, images, and attributes. We built Manticore Apparel Shop on this data.&lt;/p&gt;

&lt;p&gt;For example, you can type:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;I need black waterproof running shoes for jogging
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The system first finds suitable products, then a language model generates an answer using them as context, while the interface shows the products themselves.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Try the demo:&lt;/strong&gt; &lt;a href="https://chat.manticoresearch.com/" rel="noopener noreferrer"&gt;Manticore Apparel Shop&lt;/a&gt; generates a random product and a query that should retrieve it, then demonstrates that the same query does retrieve that product through Manticore.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvj2udfja5lxjw43kyxbq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvj2udfja5lxjw43kyxbq.png" alt="Conversational Search shows an answer and its source products in Manticore Apparel Shop" width="800" height="571"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It is important to keep the connection between the answer and the data. If the system claims that a model is suitable for rain, the user should be able to open the product and verify the source of that claim.&lt;/p&gt;

&lt;p&gt;In this approach, the language model does not replace search. It interprets its results.&lt;/p&gt;




&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;p&gt;Two main commands are used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="n"&gt;CHAT&lt;/span&gt; &lt;span class="n"&gt;MODEL&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CALL&lt;/span&gt; &lt;span class="n"&gt;CHAT&lt;/span&gt;&lt;span class="p"&gt;(...)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;First, you create a Conversational Search model and set the rules it follows.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="n"&gt;CHAT&lt;/span&gt; &lt;span class="n"&gt;MODEL&lt;/span&gt; &lt;span class="n"&gt;assistant&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'openrouter:google/gemma-4-26b-a4b-it'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;retrieval_limit&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_document_length&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;3000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;custom_prompt&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'You are a context-only shopping assistant.

Answer using only the provided context.
Do not use outside knowledge or unsupported assumptions.

Recommend only products supported by the retrieved context.
For every recommended product, briefly explain why it matches the request.

End every recommendation with the corresponding
context source ID in the format [ref:&amp;lt;id&amp;gt;].

If none of the retrieved products support the request,
say that you do not have enough information.'&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;retrieval_limit&lt;/code&gt; determines how many documents enter the context. &lt;code&gt;max_document_length&lt;/code&gt; limits the amount of text from each document.&lt;/p&gt;

&lt;p&gt;If there is too little context, the model will not see the right products. If there is too much, latency and query cost increase. Like a person, a language model does not become smarter just because it has been given everything to read.&lt;/p&gt;




&lt;p&gt;You can then run a query:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CALL&lt;/span&gt; &lt;span class="n"&gt;CHAT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="s1"&gt;'I need black waterproof running shoes for jogging'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s1"&gt;'convapparel_products'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s1"&gt;'assistant'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s1"&gt;'demo-session-001'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s1"&gt;'embedding_vector'&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You can then continue the conversation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CALL&lt;/span&gt; &lt;span class="n"&gt;CHAT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="s1"&gt;'Which of these are better for daily use?'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s1"&gt;'convapparel_products'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s1"&gt;'assistant'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s1"&gt;'demo-session-001'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s1"&gt;'embedding_vector'&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The system uses conversation history, so the user does not need to repeat the context.&lt;/p&gt;




&lt;h2&gt;
  
  
  Through the HTTP API
&lt;/h2&gt;

&lt;p&gt;Conversational Search is also available through the JSON API:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"chat"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"query"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"I need black waterproof running shoes for jogging"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"table"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"convapparel_products"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"model_name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"assistant"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"conversation_uuid"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"demo-session-001"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"vector_field"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"embedding_vector"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The request is sent to &lt;code&gt;/search&lt;/code&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  What the system returns
&lt;/h2&gt;

&lt;p&gt;The response contains:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;conversation_uuid&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;user_query&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;search_query&lt;/code&gt; — the search query generated by the system&lt;/li&gt;
&lt;li&gt;&lt;code&gt;response&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;sources&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;code&gt;search_query&lt;/code&gt; is particularly important.&lt;/p&gt;

&lt;p&gt;If the user writes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Which of these would work better in rain?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On its own, this query makes no sense without context. The system therefore forms a complete search query using the conversation history.&lt;/p&gt;

&lt;p&gt;This also simplifies debugging: you can trace the entire chain from query to answer.&lt;/p&gt;




&lt;h2&gt;
  
  
  How retrieval works
&lt;/h2&gt;

&lt;p&gt;Conversational Search uses vector search over an embedding field. The flow looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;user question
→ conversation history
→ search query
→ vector search
→ retrieved documents
→ language model
→ answer and sources
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Where search ends and conversation begins
&lt;/h2&gt;

&lt;p&gt;Full-text search works well for exact queries. Vector search works with semantic ones. Hybrid search works with their combination. Conversational Search is needed when the result must be explained, compared, or refined.&lt;/p&gt;




&lt;h2&gt;
  
  
  Quality
&lt;/h2&gt;

&lt;p&gt;We tested this on our &lt;a href="https://github.com/manticoresoftware/conversational-search-quality-becnhmark" rel="noopener noreferrer"&gt;conversational search quality benchmark&lt;/a&gt;, using 200 deterministic shopping queries from ConvApparel. The benchmark evaluates the product IDs returned as sources, rather than the wording of the generated answer.&lt;/p&gt;

&lt;p&gt;In the current run, Manticore scored &lt;strong&gt;0.3650 Hit@3&lt;/strong&gt;, &lt;strong&gt;0.4250 Hit@5&lt;/strong&gt;, &lt;strong&gt;0.5250 Hit@10&lt;/strong&gt;, and &lt;strong&gt;0.2790 MRR&lt;/strong&gt; — the best result on each of those metrics among the engines tested. The full repository includes the dataset-building rules, engine configuration, smoke tasks, and raw results.&lt;/p&gt;




&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Search in modern systems is not one algorithm but several layers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;full-text search&lt;/li&gt;
&lt;li&gt;vector search&lt;/li&gt;
&lt;li&gt;hybrid search&lt;/li&gt;
&lt;li&gt;Conversational Search&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each solves its own task. A language model does not replace search; it works on top of it.&lt;/p&gt;

&lt;p&gt;Without good search, you do not get a smart assistant; you get a very talkative consultant that barely knows its own catalog.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Want to test the approach in practice?&lt;/strong&gt; Try &lt;a href="https://chat.manticoresearch.com/" rel="noopener noreferrer"&gt;Manticore Apparel Shop&lt;/a&gt;: choose a random product, ask a question based on it, and confirm that the same product is found and suggested in the answer.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If you want to run it locally or explore the code, check the &lt;a href="https://github.com/manticoresoftware/demo-conversational-search" rel="noopener noreferrer"&gt;GitHub repository&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>database</category>
      <category>search</category>
    </item>
    <item>
      <title>UUID in Manticore: A Practical Guide</title>
      <dc:creator>Sergey Nikolaev</dc:creator>
      <pubDate>Thu, 06 Aug 2026 05:20:04 +0000</pubDate>
      <link>https://dev.to/sanikolaev/uuid-in-manticore-a-practical-guide-2foc</link>
      <guid>https://dev.to/sanikolaev/uuid-in-manticore-a-practical-guide-2foc</guid>
      <description>&lt;p&gt;In the &lt;a href="https://manticoresearch.com/blog/uuid-document-ids/" rel="noopener noreferrer"&gt;overview article&lt;/a&gt; we explained why it makes sense to use the same UUID in search and in the primary database, when one already exists. Here we will go straight to practice: create a table, run the core operations through SQL and the JSON API, and then load several documents through &lt;code&gt;/bulk&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;All examples are for Manticore Search 28.5.0 or later. The &lt;code&gt;&amp;lt;generated UUID&amp;gt;&lt;/code&gt; value in the responses means the UUID that Manticore creates when processing the request. You do not need to copy this string into the next request: substitute the actual &lt;code&gt;id&lt;/code&gt; from your own response.&lt;/p&gt;

&lt;h2&gt;
  
  
  Table for all examples
&lt;/h2&gt;

&lt;p&gt;To make UUID the document ID, first declare the field as &lt;code&gt;id uuid&lt;/code&gt;. Aside from the &lt;code&gt;id&lt;/code&gt; type, the RT table schema does not change:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;products_uuid&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="n"&gt;uuid&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;title&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;sku&lt;/span&gt; &lt;span class="n"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;price&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;DESC products_uuid&lt;/code&gt; will show that the &lt;code&gt;id&lt;/code&gt; field has the &lt;code&gt;uuid&lt;/code&gt; type. Both SQL and the HTTP API use this value as the document ID, so there is no need to copy the UUID into a string attribute.&lt;/p&gt;

&lt;h2&gt;
  
  
  Working with UUID through SQL
&lt;/h2&gt;

&lt;p&gt;First, insert a product with a prebuilt UUID:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;INSERT&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;products_uuid&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sku&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;price&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;VALUES&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="s1"&gt;'550e8400-e29b-41d4-a716-446655440000'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s1"&gt;'Mechanical keyboard'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s1"&gt;'KB-001'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="mi"&gt;149&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In SQL, the UUID must be enclosed in single quotes. You can find the document with a regular &lt;code&gt;WHERE id = '...'&lt;/code&gt; condition:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sku&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;price&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;products_uuid&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'550e8400-e29b-41d4-a716-446655440000'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Manticore can also generate the UUID. To do that, simply do not pass any &lt;code&gt;id&lt;/code&gt; value:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;INSERT&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;products_uuid&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sku&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;price&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;VALUES&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'USB microphone'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'MIC-001'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;89&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;LAST_INSERT_ID&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;LAST_INSERT_ID()&lt;/code&gt; will return the UUID created by this request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;+--------------------------------------+
| last_insert_id()                     |
+--------------------------------------+
| &amp;lt;generated UUID&amp;gt;                     |
+--------------------------------------+
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;INSERT&lt;/code&gt; for multiple documents and the &lt;code&gt;@@session.last_insert_id&lt;/code&gt; variable are covered in detail in the documentation section on &lt;a href="https://manual.manticoresearch.com/dev/Data_creation_and_modification/Adding_documents_to_a_table/Adding_documents_to_a_real-time_table" rel="noopener noreferrer"&gt;adding documents&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Let's add another document with a known ID to check &lt;code&gt;IN&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;INSERT&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;products_uuid&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sku&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;price&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;VALUES&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="s1"&gt;'550e8400-e29b-41d4-a716-446655440001'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s1"&gt;'USB-C dock'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s1"&gt;'DOCK-001'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="mi"&gt;119&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sku&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;price&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;products_uuid&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="k"&gt;IN&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="s1"&gt;'550e8400-e29b-41d4-a716-446655440000'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s1"&gt;'550e8400-e29b-41d4-a716-446655440001'&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You can change attributes with a regular &lt;code&gt;UPDATE&lt;/code&gt;. The &lt;code&gt;id&lt;/code&gt; itself stays the same:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;UPDATE&lt;/span&gt; &lt;span class="n"&gt;products_uuid&lt;/span&gt;
&lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;price&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;139&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'550e8400-e29b-41d4-a716-446655440000'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;An &lt;code&gt;INSERT&lt;/code&gt; with an already existing UUID will not overwrite the document. Manticore will return a duplicate error:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;INSERT&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;products_uuid&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sku&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;price&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;VALUES&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="s1"&gt;'550e8400-e29b-41d4-a716-446655440000'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s1"&gt;'Duplicate keyboard'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s1"&gt;'KB-DUP'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="mi"&gt;1&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;To create a new version of the document with the same UUID, use &lt;code&gt;REPLACE&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;REPLACE&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;products_uuid&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sku&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;price&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;VALUES&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="s1"&gt;'550e8400-e29b-41d4-a716-446655440000'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s1"&gt;'Mechanical keyboard, revised'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s1"&gt;'KB-001'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="mi"&gt;129&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;To change only the price, &lt;code&gt;UPDATE&lt;/code&gt; is enough. Full-text fields and columnar attributes require &lt;code&gt;REPLACE&lt;/code&gt;: it marks the old version of the document with the same ID as deleted and writes the new one. If that ID does not exist yet, Manticore simply adds the document. More details are in the documentation: &lt;a href="https://manual.manticoresearch.com/dev/Data_creation_and_modification/Updating_documents/UPDATE" rel="noopener noreferrer"&gt;UPDATE&lt;/a&gt; and &lt;a href="https://manual.manticoresearch.com/dev/Data_creation_and_modification/Updating_documents/REPLACE" rel="noopener noreferrer"&gt;REPLACE&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Delete the document by the same UUID:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;DELETE&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;products_uuid&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'550e8400-e29b-41d4-a716-446655440001'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Let's check the current state of the document whose UUID we set on insert:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sku&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;price&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;products_uuid&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'550e8400-e29b-41d4-a716-446655440000'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The SQL client sends and receives a UUID as a string. Pass it as a string parameter, and read the &lt;code&gt;id&lt;/code&gt; from &lt;code&gt;SELECT&lt;/code&gt; as a string. There is no need to convert it to a number or &lt;code&gt;BINARY(16)&lt;/code&gt;. Code written for a numeric document ID will need to be adjusted.&lt;/p&gt;

&lt;h2&gt;
  
  
  The same operations through the JSON API
&lt;/h2&gt;

&lt;p&gt;When writing through the JSON API, the &lt;code&gt;id&lt;/code&gt; field is passed alongside &lt;code&gt;table&lt;/code&gt;, not inside &lt;code&gt;doc&lt;/code&gt;. In the first example, we use an uppercase UUID:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sS&lt;/span&gt; http://localhost:9308/insert &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Content-Type: application/json'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "table": "products_uuid",
    "id": "AAAAAAAA-AAAA-4AAA-8AAA-AAAAAAAAAAAA",
    "doc": {
      "title": "Wireless keyboard",
      "sku": "JSON-KB-001",
      "price": 159
    }
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Manticore accepts UUID in uppercase, stores it in lowercase, and returns it lowercased in the response:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"table"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"products_uuid"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"aaaaaaaa-aaaa-4aaa-8aaa-aaaaaaaaaaaa"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"created"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"result"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"created"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;201&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For automatic generation, remove the &lt;code&gt;id&lt;/code&gt; field entirely:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sS&lt;/span&gt; http://localhost:9308/insert &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Content-Type: application/json'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "table": "products_uuid",
    "doc": {
      "title": "Portable speaker",
      "sku": "JSON-SPK-001",
      "price": 79
    }
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The response contains the ID that should be saved for subsequent operations:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"table"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"products_uuid"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"&amp;lt;generated UUID&amp;gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"created"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"result"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"created"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;201&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In &lt;code&gt;/search&lt;/code&gt;, you can use UUID in the &lt;code&gt;equals&lt;/code&gt; filter. If you request &lt;code&gt;id&lt;/code&gt; in &lt;code&gt;_source&lt;/code&gt;, the result will contain the same UUID in both &lt;code&gt;_id&lt;/code&gt; and &lt;code&gt;_source.id&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sS&lt;/span&gt; http://localhost:9308/search &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Content-Type: application/json'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "table": "products_uuid",
    "query": {
      "equals": {
        "id": "aaaaaaaa-aaaa-4aaa-8aaa-aaaaaaaaaaaa"
      }
    },
    "_source": ["id", "title", "sku", "price"]
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"timed_out"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"hits"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"total"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"total_relation"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"eq"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"hits"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"aaaaaaaa-aaaa-4aaa-8aaa-aaaaaaaaaaaa"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"_score"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"_source"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"aaaaaaaa-aaaa-4aaa-8aaa-aaaaaaaaaaaa"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"title"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Wireless keyboard"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"sku"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"JSON-KB-001"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"price"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;159&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;_id&lt;/code&gt; is part of the search result metadata, while &lt;code&gt;_source.id&lt;/code&gt; is the document field. For a UUID table, they contain the same string.&lt;/p&gt;

&lt;p&gt;Now let's call the remaining endpoints for modifying data one by one. &lt;code&gt;UPDATE&lt;/code&gt; changes only the price:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sS&lt;/span&gt; http://localhost:9308/update &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Content-Type: application/json'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "table": "products_uuid",
    "id": "aaaaaaaa-aaaa-4aaa-8aaa-aaaaaaaaaaaa",
    "doc": {"price": 149}
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;To replace the document completely, call &lt;code&gt;/replace&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sS&lt;/span&gt; http://localhost:9308/replace &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Content-Type: application/json'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "table": "products_uuid",
    "id": "aaaaaaaa-aaaa-4aaa-8aaa-aaaaaaaaaaaa",
    "doc": {
      "title": "Wireless keyboard, revised",
      "sku": "JSON-KB-001",
      "price": 139
    }
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now delete the replaced document:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sS&lt;/span&gt; http://localhost:9308/delete &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Content-Type: application/json'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "table": "products_uuid",
    "id": "aaaaaaaa-aaaa-4aaa-8aaa-aaaaaaaaaaaa"
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After &lt;code&gt;/delete&lt;/code&gt;, there is no document with this UUID left in the table. A &lt;code&gt;/search&lt;/code&gt; by the same ID will return &lt;code&gt;total: 0&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sS&lt;/span&gt; http://localhost:9308/search &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Content-Type: application/json'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "table": "products_uuid",
    "query": {
      "equals": {
        "id": "aaaaaaaa-aaaa-4aaa-8aaa-aaaaaaaaaaaa"
      }
    }
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"timed_out"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"hits"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"total"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"total_relation"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"eq"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"hits"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Examples of requests for deleting by ID and by condition are collected in the documentation section on &lt;a href="https://manual.manticoresearch.com/dev/Data_creation_and_modification/Deleting_documents" rel="noopener noreferrer"&gt;deleting documents&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where to generate UUID
&lt;/h2&gt;

&lt;p&gt;If the UUID is already issued by the primary database, just pass it to Manticore as &lt;code&gt;id&lt;/code&gt;. If the request is repeated, the UUID will stay the same. &lt;code&gt;INSERT&lt;/code&gt; will report a duplicate, and &lt;code&gt;REPLACE&lt;/code&gt; will write a new version of the document under the same ID. The same rule applies to &lt;code&gt;insert&lt;/code&gt; and &lt;code&gt;replace&lt;/code&gt; in &lt;code&gt;/bulk&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Manticore can also generate the UUID itself: just do not pass &lt;code&gt;id&lt;/code&gt;. But resending such a request will create another document, so for automatic retries it is better to set the UUID explicitly.&lt;/p&gt;

&lt;p&gt;For an explicit ID, any UUID version from v1 to v8 is suitable. For automatic generation, Manticore uses its own UUIDv8 structure, but it does not convert UUIDs received from the client into it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Batch loading through &lt;code&gt;/bulk&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;/bulk&lt;/code&gt; accepts NDJSON: each line contains a separate operation. The UUID is passed in the &lt;code&gt;id&lt;/code&gt; field, just like in the other JSON API requests:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;POST /bulk
Content-Type: application/x-ndjson

{"insert":{"table":"products_uuid","id":"bbbbbbbb-bbbb-4bbb-8bbb-bbbbbbbbbbbb","doc":{"title":"USB hub","sku":"BULK-HUB-001","price":49}}}
{"insert":{"table":"products_uuid","id":"bbbbbbbb-bbbb-4bbb-8bbb-bbbbbbbbbbbc","doc":{"title":"Laptop stand","sku":"BULK-STAND-001","price":39}}}
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After the last data line, a trailing newline is required. With &lt;code&gt;curl&lt;/code&gt;, it is convenient to pass the body through &lt;code&gt;--data-binary&lt;/code&gt; so it does not strip newlines.&lt;/p&gt;

&lt;p&gt;Both operations belong to one table, so Manticore executes them in a single transaction. The shortened response shows how many documents were added and how the entire batch finished:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"items"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"bulk"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"created"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;201&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"current_line"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"skipped_lines"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"errors"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On an error, &lt;code&gt;current_line&lt;/code&gt; shows the line where processing stopped, and &lt;code&gt;skipped_lines&lt;/code&gt; shows the number of skipped lines. If an empty line or a table switch split the request into multiple transactions, Manticore will not roll back the ones that already completed.&lt;/p&gt;

&lt;p&gt;When reloading a document, it is important to choose the right operation. If the UUID already exists, &lt;code&gt;insert&lt;/code&gt; will fail with a duplicate error. &lt;code&gt;replace&lt;/code&gt; will write the document again, and if no such ID exists, it will add a new document.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;id&lt;/code&gt; field in &lt;code&gt;/bulk&lt;/code&gt; can also be omitted, and Manticore will generate a UUID. But the &lt;code&gt;/bulk&lt;/code&gt; response contains only the overall transaction result, without separate results for each inserted row. If you need to keep the ID of each document, it is more convenient to generate the UUID before the batch request or add documents one by one through &lt;code&gt;/insert&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Validation and common errors
&lt;/h2&gt;

&lt;p&gt;Manticore validates the UUID before adding the document. SQL, the JSON API, and &lt;code&gt;/bulk&lt;/code&gt; return errors differently, so do not bind your code to the exact wording of the message.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What you sent&lt;/th&gt;
&lt;th&gt;What Manticore will do&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;550e8400-e29b-41d4-a716-446655440000&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Add the document&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;AAAAAAAA-AAAA-4AAA-8AAA-AAAAAAAAAAAA&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Add the document, store the UUID, and return it in lowercase&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A string without hyphens or with an invalid character&lt;/td&gt;
&lt;td&gt;Return an error because the UUID format is invalid&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;00000000-0000-0000-0000-000000000000&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Return an error: zero UUID cannot be used&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A number, including &lt;code&gt;0&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Return an error: &lt;code&gt;id&lt;/code&gt; must be a string&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A UUID with a version outside the 1-8 range or an RFC-invalid &lt;code&gt;variant&lt;/code&gt; value&lt;/td&gt;
&lt;td&gt;Return a validation error&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;INSERT&lt;/code&gt; or &lt;code&gt;insert&lt;/code&gt; in &lt;code&gt;/bulk&lt;/code&gt; with an already existing ID&lt;/td&gt;
&lt;td&gt;Return a duplicate error&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A few points:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The canonical form consists of 36 characters split into &lt;code&gt;8-4-4-4-12&lt;/code&gt; groups. UUID versions from 1 to 8 are valid; in the RFC &lt;code&gt;variant&lt;/code&gt; position, one of these characters must appear: &lt;code&gt;8&lt;/code&gt;, &lt;code&gt;9&lt;/code&gt;, &lt;code&gt;a&lt;/code&gt;, or &lt;code&gt;b&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Manticore validates only the ID format. Correct generation is the application's responsibility: for UUIDv4, randomness matters, and for UUIDv7, the time component and compliance with the rules matter. Manticore does not convert passed v4 and v7 values into v8.&lt;/li&gt;
&lt;li&gt;In a table with UUID IDs, &lt;code&gt;id = 0&lt;/code&gt; does not trigger automatic generation unlike numeric IDs: the request will fail. But if you do not pass &lt;code&gt;id&lt;/code&gt;, Manticore will generate its own UUIDv8 structure and return it in the response. Keep in mind that it can be used as a document ID, but not as, for example, an access token.&lt;/li&gt;
&lt;li&gt;You can validate the UUID with a standard library on input, but keep in mind that Manticore still performs its own validation.&lt;/li&gt;
&lt;li&gt;In the &lt;code&gt;/bulk&lt;/code&gt; response, check &lt;code&gt;errors&lt;/code&gt;, &lt;code&gt;current_line&lt;/code&gt;, and &lt;code&gt;skipped_lines&lt;/code&gt;: they show where processing stopped and which part of the batch could not be committed. In SQL batch requests, the entire request is rolled back if an error occurs.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Columnar RT and replication
&lt;/h2&gt;

&lt;p&gt;UUID can also be used as the document ID in RT tables with columnar storage. Only the engine declaration changes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;products_uuid_columnar&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="n"&gt;uuid&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;title&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;sku&lt;/span&gt; &lt;span class="n"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;price&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;engine&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'columnar'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For regular and columnar RT tables, &lt;code&gt;INSERT&lt;/code&gt;, &lt;code&gt;REPLACE&lt;/code&gt;, and &lt;code&gt;DELETE&lt;/code&gt;, exact-match searches, and &lt;code&gt;IN&lt;/code&gt; conditions are written the same way. The client code does not depend on the table engine. But &lt;code&gt;UPDATE&lt;/code&gt; does not change columnar attributes, so, for example, the price in such a table can only be updated together with the whole document through &lt;code&gt;REPLACE&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Replication also supports UUID. Suppose the &lt;code&gt;catalog&lt;/code&gt; cluster has already been created, the nodes have joined, and the local &lt;code&gt;products_uuid&lt;/code&gt; table exists. Let's add it to the cluster:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;CLUSTER&lt;/span&gt; &lt;span class="k"&gt;catalog&lt;/span&gt; &lt;span class="k"&gt;ADD&lt;/span&gt; &lt;span class="n"&gt;products_uuid&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In SQL, a colon goes between the cluster name and the table name:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;INSERT&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="k"&gt;catalog&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;products_uuid&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sku&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;price&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;VALUES&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="s1"&gt;'550e8400-e29b-41d4-a716-446655441000'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s1"&gt;'Replicated keyboard'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s1"&gt;'REPL-KB-001'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="mi"&gt;169&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sku&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;price&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="k"&gt;catalog&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;products_uuid&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'550e8400-e29b-41d4-a716-446655441000'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In the JSON API, the table name is passed unchanged, while the cluster goes in a separate field. For example, the following operation will update a document previously added through SQL:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sS&lt;/span&gt; http://localhost:9308/update &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Content-Type: application/json'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "cluster": "catalog",
    "table": "products_uuid",
    "id": "550e8400-e29b-41d4-a716-446655441000",
    "doc": {"price": 159}
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After replication, the document keeps the same UUID on all nodes. You can run exact-match searches, &lt;code&gt;UPDATE&lt;/code&gt;, &lt;code&gt;REPLACE&lt;/code&gt;, and &lt;code&gt;DELETE&lt;/code&gt; by it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to consider before rollout
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The &lt;code&gt;uuid&lt;/code&gt; type can be assigned only to the &lt;code&gt;id&lt;/code&gt; field. It is supported in RT tables, including those with columnar storage and replication, but not in plain, percolate/PQ, or shard tables.&lt;/li&gt;
&lt;li&gt;An existing table cannot be switched from a numeric ID to a UUID or back with &lt;code&gt;ALTER TABLE&lt;/code&gt;. This decision has to be made when creating the new schema.&lt;/li&gt;
&lt;li&gt;A UUID can be used in &lt;code&gt;=&lt;/code&gt; and &lt;code&gt;IN&lt;/code&gt; conditions. The &lt;code&gt;&amp;lt;&lt;/code&gt;, &lt;code&gt;&amp;lt;=&lt;/code&gt;, &lt;code&gt;&amp;gt;&lt;/code&gt;, and &lt;code&gt;&amp;gt;=&lt;/code&gt; ranges, as well as arithmetic on &lt;code&gt;id&lt;/code&gt;, are not supported.&lt;/li&gt;
&lt;li&gt;UUIDv7 contains time, but filtering &lt;code&gt;id&lt;/code&gt; by range is still not allowed. For date-based selection, add an attribute such as &lt;code&gt;created_at timestamp&lt;/code&gt; and filter by it.&lt;/li&gt;
&lt;li&gt;The document ID cannot be changed with &lt;code&gt;UPDATE&lt;/code&gt;. To give an object a different UUID, you will need to create a document with a new ID and delete the old one separately.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The behavior of SQL sessions and the full list of limitations are described in the documentation: &lt;a href="https://manual.manticoresearch.com/dev/Creating_a_table/Data_types#UUID-document-IDs" rel="noopener noreferrer"&gt;UUID document IDs&lt;/a&gt;. Support for UUID as a document ID appeared in &lt;a href="https://manticoresearch.com/blog/manticore-search-28-6-6/" rel="noopener noreferrer"&gt;Manticore Search 28.6.6&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>database</category>
      <category>search</category>
      <category>sql</category>
      <category>tutorial</category>
    </item>
  </channel>
</rss>
