An embedding model doesn't generate text — it converts text into a list of numbers that represents what that text means. Two pieces of writing about similar topics end up with similar vectors even if they don't share a single word, which is the whole trick behind semantic search: you can find "things that mean something like this," not just "things that contain this exact word." The detail that actually surprised me: the vector is always the same fixed length no matter how long the source text is. A one-sentence note and a ten-page article both come out as the same-sized list of numbers — a fingerprint, basically, summarizing something much bigger into a fixed shape.
I built this up one command at a time instead of writing the whole script upfront, mostly because I wanted to actually see what each piece was doing before gluing them together.
Pulled the embedding model first:
ollama pull nomic-embed-text
Then sanity-checked it worked before writing any Python at all — Ollama exposes embeddings over a local HTTP API, so a plain curl gets you a real answer with zero code:
curl http://localhost:11434/api/embeddings -d '{"model": "nomic-embed-text", "prompt": "hello world"}'
Comes back with a JSON object holding one embedding field — a list of 768 numbers. That's genuinely the whole idea in one command: text in, fixed-length list of numbers out.
Installed the Python pieces:
pip3 install ollama chromadb requests beautifulsoup4 --break-system-packages
Embedded one local file, a few lines at a time in a Python shell rather than running a full script blind:
import ollama
text = open("02-oc-cli-mentor-system-prompt.md").read()
response = ollama.embeddings(model="nomic-embed-text", prompt=text)
len(response["embedding"]) # → 768
Hit ModuleNotFoundError: No module named 'ollama' on the first try — I'd skipped the install step above without noticing. Even a five-step walkthrough apparently has room to trip over your own feet.
Stored it in Chroma:
import chromadb
client = chromadb.PersistentClient(path="./chroma_db")
collection = client.get_or_create_collection(name="today_i_ran_notes")
collection.upsert(
ids=["02-oc-cli-mentor-system-prompt.md"],
embeddings=[response["embedding"]],
documents=[text],
)
collection.count() # → 1
Added a longer file next, and hit a real wall. Repeating the same embed call against a longer draft failed differently:
ollama._types.ResponseError: the input length exceeds the context length
Turns out nomic-embed-text has a 2048-token context window, and the longer draft — 13,985 characters — blew straight past it. Quick, honestly not great fix: just truncate.
text = text[:6000]
Kept me moving, but this is a real problem, not a footnote — more on that below.
Added a live URL as a third source. Pulling in a web page instead of a local file needs one extra step: fetch the page, then strip out the nav bars and scripts and footers so only the actual article text gets embedded.
import requests
from bs4 import BeautifulSoup
response = requests.get("https://pipelineandprompts.com/posts/03-1b-vs-3b-memory-comparison/",
headers={"User-Agent": "Mozilla/5.0"})
soup = BeautifulSoup(response.text, "html.parser")
for tag in soup(["script", "style", "nav", "footer", "header", "aside"]):
tag.decompose()
article_text = soup.get_text(separator="\n", strip=True)
len(article_text) # → 3359
Then fed that into the same embedding call from before.
Ended up with three sources embedded:
| Source | Type | Characters sent | Vector dimensions |
|---|---|---|---|
cloud-without-chaos-01.md |
local file | 6,000 (truncated from 13,985) | 768 |
02-oc-cli-mentor-system-prompt.md |
local file | 2,941 | 768 |
| Entry 03 (live URL) | web article | 3,359 | 768 |
Two things stood out. Every vector came back at exactly 768 dimensions no matter the source length — confirms the fixed-size fingerprint idea wasn't just something I read somewhere. The more important thing: that truncation fix meant 57% of the longest draft's content never made it into its embedding at all — 7,985 of 13,985 characters, just gone. That's not a rounding error, that's most of an article being invisible to any future search against it. This is exactly why real RAG pipelines chunk long documents into overlapping pieces instead of truncating — chunking keeps everything searchable, truncation just quietly throws stuff away and doesn't tell you.
The URL fetch worked cleanly, for what it's worth — 3,359 characters extracted is close to the article's real body length, so the boilerplate-stripping did its job without me having to fight it.
The pipeline works end to end — files and live URLs, embedded into something queryable — but truncation is a real bug here, not a technical footnote, the moment you're dealing with anything longer than a short note. Chunking is next. Also still completely untested: whether querying this thing actually surfaces the right document for a real question. That's the actual point of building it, and I haven't gotten there yet.
Top comments (0)