I built PixaFind for the search you do when you remember a file's contents but have forgotten its name. You might remember a cat sitting in sunlight, a whiteboard from a meeting, or a paragraph about a budget. None of those memories give you a path to type into Explorer.
On Windows, I use a Tauri 2 shell, a Rust backend, ONNX Runtime inference, and local SQLite vector storage. SigLIP2 handles text-to-image search. DINOv3 handles image-to-image similarity. For documents, I combine dense embeddings with a BM25 keyword index.
The engineering cost sits in the indexing pipeline: decoding files, preparing model inputs, running inference, and keeping the index consistent as users edit their folders.
Put the inference boundary on the user's machine
I keep file access, inference, and retrieval in the Rust backend. Users interact with a floating search window through a global hotkey, Ctrl+Space, without switching away from their current application.
The native Tauri shell weighs about 5 MB. That figure describes the executable, not the full installation. Users also need model weights, ONNX Runtime dependencies, and index storage. The release README describes a one-time model download of roughly 500 MB.
Tauri uses the system webview on Windows. I avoid shipping a separate Chromium distribution, but I still have to account for webview memory, model memory, and GPU allocations. A small executable does not imply a 5 MB runtime footprint.
This diagram shows the architecture at the component level, rather than private source-code interfaces:
flowchart LR
F["Selected folders"] --> W["Filesystem watcher"]
W --> C["Rust: mtime and SHA-256 change detection"]
C --> P["Decode images or extract document text"]
P --> O["ONNX Runtime: DirectML, CUDA, or CPU"]
O --> V["Vector DB: local SQLite storage"]
P --> B["BM25 keyword index"]
V --> R["Rust retrieval and ranking"]
B --> R
H["Global hotkey"] --> U["Spotlight UI: Tauri 2"]
U --> Q["Query or selected image"]
Q --> O
R --> U
I separate background indexing from interactive retrieval because the two workloads compete for resources. During indexing, a user may still want to search the files that have finished processing. An implementation needs bounded work queues and an explicit policy for sharing GPU time; launching an inference task for every filesystem event would exhaust memory before it improved throughput.
Use SigLIP2 for descriptions, DINOv3 for visual neighbors
For a query such as cat in a sunbeam, I need an image representation that shares a space with a text representation. SigLIP2 supplies that alignment. During indexing, I encode an image. At search time, I encode the user's description and compare it with the stored image embeddings.
The user does not need to tag the image or include the word cat in its filename. Zero-shot matching means I can search with a description without training a classifier for each phrase.
For normalized embeddings, cosine similarity reduces to a dot product:
query text -> SigLIP2 text encoder -> q
photo -> SigLIP2 image encoder -> v
score(q, v) = sum(q[i] * v[i])
A description such as blurry photo of whiteboard asks for visual meaning. It does not guarantee that I can retrieve a sentence written on that whiteboard. Users still depend on the image resolution and the model's ability to recognize the scene. Exact text inside an image requires an OCR path; semantic image matching alone does not establish one.
When a user selects a photo and asks for similar images, I use DINOv3. With that route, I compare visual embeddings from the selected image against visual embeddings from the library. The user can look for similar compositions or objects without translating them into a sentence.
I keep the two embedding spaces separate. A DINOv3 vector and a SigLIP2 text vector do not acquire compatible meaning because both contain floating-point numbers. Each search route needs the matching model, preprocessing, and vector collection.
flowchart TD
T["Text: cat in a sunbeam"] --> ST["SigLIP2 text encoder"]
I["Indexed image"] --> SI["SigLIP2 image encoder"]
ST --> S["Compare within SigLIP2 space"]
SI --> S
X["Selected image"] --> DX["DINOv3 encoder"]
I --> DI["DINOv3 encoder"]
DX --> D["Compare within DINOv3 space"]
DI --> D
S --> A["Rank image results"]
D --> A
Run ONNX inference without guessing the model contract
I use ONNX Runtime to execute the exported models on Windows. PixaFind supports DirectML and CUDA acceleration, with a CPU fallback. DirectML covers compatible DirectX 12 GPUs across vendors; CUDA targets supported NVIDIA hardware.
I still need to ship a runtime that contains the selected execution provider. Enabling a Rust Cargo feature does not install CUDA libraries or give an ONNX Runtime binary DirectML support. I also need to check the exported graph's operators, tensor shapes, and provider compatibility.
The following example uses the documented API from ort 2.0.0-rc.10. It illustrates a DirectML image-embedding inference boundary, not a copy of PixaFind's private implementation. The caller supplies the actual model path, tensor names, spatial dimensions, and preprocessed pixels from its export contract.
use std::path::Path;
use ort::{
execution_providers::DirectMLExecutionProvider,
session::Session,
value::TensorRef,
};
fn open_image_encoder(model_path: &Path) -> ort::Result<Session> {
Session::builder()?
.with_parallel_execution(false)?
.with_memory_pattern(false)?
.with_execution_providers([
DirectMLExecutionProvider::default().build(),
])?
.commit_from_file(model_path)
}
fn infer_embedding(
session: &mut Session,
input_name: &str,
output_name: &str,
height: usize,
width: usize,
pixels_nchw: &[f32],
destination: &mut Vec<f32>,
) -> ort::Result<()> {
// The tensor view borrows the caller's preprocessed image buffer.
let input = TensorRef::from_array_view((
[1usize, 3, height, width],
pixels_nchw,
))?;
let outputs = session.run(ort::inputs![input_name => input])?;
let (_shape, embedding) =
outputs[output_name].try_extract_tensor::<f32>()?;
// Preserve the result after the runtime output values leave scope.
destination.clear();
destination.extend_from_slice(embedding);
Ok(())
}
I reuse the session across images instead of loading model weights for each forward pass. I borrow the input buffer and copy the output into caller-owned storage because I need that embedding after the runtime releases its output values. The caller can reuse the destination allocation across files.
The caller must decode the image, apply the export's resize or crop rule, normalize channels with the correct constants, and produce the expected NCHW layout. It must also verify that the chosen output contains a pooled image embedding rather than logits or a sequence of patch features. Changing any of those choices can change retrieval quality without triggering a tensor-type error.
For SigLIP2 text inference, I need a separate input contract: the matching tokenizer, token IDs, padding policy, and any required mask tensors. I cannot substitute image preprocessing or an unrelated tokenizer.
Microsoft's DirectML execution-provider documentation requires sequential graph execution and disabled memory-pattern optimization. It also prohibits concurrent Run calls on the same DirectML session. That constraint affects worker design: I serialize access to a shared session, or provision separate sessions and accept their additional memory cost.
The documentation now describes DirectML as being in sustained engineering, with new Windows deployment development moving toward WinML. For developers choosing a runtime today, that distinction matters even when an existing DirectML deployment continues to work.
Keep exact document terms alongside dense embeddings
For document search, I combine dense vector retrieval with BM25. The release README documents approximately 30 document extensions, including .pdf, .docx, .pptx, .md, .txt, and source-code formats. The product website advertises a broader 76-plus-format count. I use the release README's narrower document scope here rather than treating those two counts as interchangeable.
A developer searching for vacation policy may want a paragraph about annual leave. Dense embeddings can retrieve that semantic relationship. Another developer searching for ERR_CONNECTION_RESET, a function name, or an invoice identifier needs exact terms. BM25 covers that lexical route.
For a hybrid implementation, I would use this retrieval sequence:
Extracted document text
|
+-> bounded text chunks -> dense embeddings -> semantic candidates
|
+-> tokenized text -> BM25 index -> keyword candidates
|
merge, deduplicate, rank by file
This sketch describes a design pattern, not a claim about PixaFind's private chunk size or fusion formula. The public release repository documents hybrid search without publishing those internals.
Developers need a ranking policy because BM25 scores and vector similarities have different scales. Adding the raw values can let one route dominate by numerical range. Rank-based fusion offers one option: assign a contribution according to a candidate's position in each list, then deduplicate document chunks before presenting file results. Tuning that policy requires queries from the target library, including both paraphrases and exact identifiers.
Text extraction also sets a ceiling on recall. A scanned PDF may contain pixels without a text layer. Office documents can include tables whose reading order affects the extracted text. Source-code tokenization can split identifiers that users expect to search as a unit. I would inspect extraction output before blaming the embedding model for a missing result.
Avoid paying for the same forward pass twice
I use modification-time checks and SHA-256 content hashing to avoid redundant inference on unchanged files. Reading metadata costs less than decoding an image and running a vision encoder. Hashing costs a full read of the file, so I do not treat it as free.
The conceptual decision tree looks like this:
Filesystem event
|
v
Compare stored metadata with current metadata
|
+-> unchanged: retain the existing index entry
|
v
Compute SHA-256 for a changed candidate
|
+-> same content digest: update metadata, retain embeddings
|
v
Decode or extract -> run inference -> replace index entries
Modification times provide a cheap filter, not proof of identical content. Users can restore timestamps, and some workflows preserve them while replacing files. A correctness-focused implementation needs reconciliation or content verification for those cases. It also needs to avoid committing an embedding for bytes that changed midway through the read.
For index maintenance, I would update file metadata and its associated vectors in one SQLite transaction. Deletions need removal from both vector and keyword search. A rename should preserve the relationship between the searchable content and its current path rather than leave an obsolete result that the user cannot open.
The content digest answers whether the file changed. It does not answer whether its embedding contract changed. After upgrading a model, tokenizer, or image preprocessing rule, developers need to invalidate the affected embeddings even when the source bytes remain identical. Mixing model revisions inside one vector collection produces rankings that users cannot interpret.
Budget for the index as well as the models
Local SQLite storage keeps vectors and file metadata on the user's disk. I do not need a remote vector service for this architecture, but I still need a retrieval strategy that fits the library size. Storing vectors in SQLite alone does not establish whether an implementation uses an exact scan or an approximate nearest-neighbor index.
For float32 storage, I can estimate the vector payload before accounting for database pages and metadata:
payload bytes = item count * dimensions * 4
Illustrative sizing, not PixaFind's model dimensions:
100,000 items * 768 dimensions * 4 bytes = 307,200,000 bytes
about 293 MiB
If I store a second image embedding for visual similarity, I add its own dimensionality and item count to that calculation. Document chunks can create several vectors per file. GPU inference reduces compute time; it does not reduce those disk costs.
I would measure decode time, inference time, database write time, and interactive query latency as separate stages. Without hardware, model export, batch size, and library composition, a single indexing-speed number tells a developer little. This article makes no benchmark claim for those stages.
Draw the offline boundary around setup and search
PixaFind performs indexing and search on the machine, stores the index locally, and sends no telemetry. Users do not upload filenames, file contents, thumbnails, or embeddings to a search service.
Installation has a separate lifecycle. The release README describes first-launch model downloads and an in-app updater. Users therefore need network access for normal initial provisioning, even though subsequent search needs none. I would not describe that bootstrap process as zero network activity.
For an air-gapped deployment, an operator needs to provision the installer, runtime dependencies, and complete model assets before disconnecting the machine. Offline use assumes those assets already exist locally. The public README does not document a dedicated offline model-import command, so I cannot offer one here.
Local storage also requires a threat model. An embedding index can expose information about the source library to someone with access to that Windows account or disk. Keeping the computation local removes the remote search-service boundary; users still need filesystem permissions, disk encryption where appropriate, and a policy for deleting index data.
Try the search routes on your own folders
Download the Windows installer from PixaFind Releases, complete model provisioning, and choose a folder you want to index. The public repository hosts installers and update artifacts. PixaFind is commercial software, and I have not published its application source there.
To evaluate retrieval, choose queries with different demands:
| Route | Example | Check |
|---|---|---|
| Text-to-image | cat in a sunbeam |
Does the result match the remembered scene? |
| Text-to-image | blurry photo of whiteboard |
Can you find the photo without knowing its name? |
| Image-to-image | Select a photo and request similar images | Do the neighbors share useful visual features? |
| Document semantics | vacation policy |
Can you find a relevant passage phrased as annual leave? |
| Document keywords | An identifier you know exists in a file | Does hybrid ranking retain the exact match? |
Then edit one indexed file and leave the others untouched. That exercise exposes a different part of the system: freshness, deletion handling, and the cost of incremental work. After provisioning, repeat your searches without network connectivity to evaluate the offline path.
I chose this architecture so users can search from descriptions while keeping their files and query processing on Windows. Its maintenance work remains concrete: compatible model exports, correct preprocessing, provider-specific session rules, and an index that tracks the files users have today.
Written by Ishan Naik. Product: PixaFind. Downloads and public documentation: pixafind-releases.

Top comments (0)