DEV Community

Cover image for Is It Still RAG If There's No Retrieval?
Jason Agostoni
Jason Agostoni

Posted on Originally published at jason.agostoni.net

Is It Still RAG If There's No Retrieval?

I've been shipping software for 28 years, 21 of them as a consultant working across over 100 different engagements. Past a certain point, a resume stops being a record and becomes a burdensome editing problem. Two pages can barely hold a half-dozen engagements, selected by hand for someone to read cold before I even get a chance to chat with them. Everything else is real experience that somebody, somewhere, might want to know.

It bothered me that the odds are decent the one engagement a particular hiring manager is looking for didn't make the cut. Maybe the point-of-sale integration story that would have mattered to a retail-tech team got bumped for something that better fit the job description.

So I built a (sigh, another) chatbot. As I put myself into the job market, an idea that had been sitting there since my last product shipped resurfaced: put my career history, skills, hobbies, and interests behind a single interface anyone can query, and ship it as an actual product rather than a demo. Two birds, one side quest.

What I ended up with is a small Go and HTMX app with one curated Markdown file as its entire knowledge base: no embeddings, no vector store, no agent loop anywhere. It's live; you'll have to visit my LinkedIn profile or snag a copy of my resume for the link. Whether that even counts as RAG is a fair question, which I hope to answer in this article. What follows is how I constrained it, what I built, and where I had the most fun.

The design brief

I set the constraints before writing any code as they would help shape the implementation and technology choices.

Nothing I'm not allowed to say. Twenty-one years of consulting means NDAs, confidentiality agreements, and client names all through my records. The source data had to be treated as a leak risk before it could become knowledge. Client names became anonymized descriptors ("a national grocery chain" instead of the real one), contact details and commercial terms were stripped, and anything under NDA was excluded outright. Then review passes, plural: an agentic review, a mechanical keyword sweep for every client name I could remember, a manual read of the whole file, and a regeneration. The entire knowledge boundary is one curated Markdown file. If it isn't in the file, the bot can't say it.

Me, minus most of me. No address, no phone, no email. The only personal details in the bundle are the ones I put there on purpose: location, availability, and hobbies and interests for the culture-fit questions. LinkedIn is the contact route, full stop. My personal information is in the hands of enough shady data brokers already.

Never invent. The knowledge boundary dictates what the bot can answer; this rule sets how it answers: only facts and metrics that appear explicitly in the source data. No rounding up engagement counts, no inferring a skill from an adjacent one, no helpful guesses. When a question falls outside the bundle, the bot gently declines to answer. A chatbot that confabulates my experience is worse than the resume problem I started with.

No opinions. The bot is a little more than a fact retriever. Asked "is Jason a fit for this role?", it answers with the facts the bundle provides and suggests that would be a great thing to bring up in an interview. It doesn't editorialize about my strengths or volunteer a verdict. Judging fit is the hiring team's job; the bot's job is making sure they're judging complete information.

It is not an agent

A key design decision came next: the bot has no tools, no ReAct loop, no reasoning cycle of any kind. It is just a query tool. You ask a question, it answers from the bundle, simple and done.

An "agent" didn't make sense here, so everything shifted to pre-processing and bundle build time. Look things up? The entire knowledge base is in the prompt. Compute years of .NET experience? All aggregates were precomputed. Take an action? There is nothing for it to do: the product is a conversation.

Stripping the agent loop streamlined things. There is no tool-calling code to write, no schema to maintain, no loop to debug when the model decides to call the wrong thing three times in a row. Requests are mini-sessions and responses are fast, and the failures are boring. It shrinks the security surface too: an agent that can call tools and take actions gives an attacker something to reach for, and this bot can only speak.

To "RAG" or not to RAG

Retrieval-augmented generation (RAG) is the standard answer when you want a model to answer from your data: chunk the documents, embed them, drop them in a vector store, and fetch the top matches for each question. My entire career knowledge base fits in a modern context window with room to spare, so at runtime there was nothing to retrieve. Every chunk is always "retrieved."

Instead of abandoning the acronym, I re-staged its three phases.

Retrieval happened offline, long before any user showed up, and before I even started the bot. It was a mining and cleansing process: extracting the history from client dossiers and notes, applying the curation gate, and staging the result as a single Markdown file. The retrieval question, what the bot should know, got answered once, reviewed and edited by hand, carefully.

Augmentation happens at the start of every chat session, as part of the conversation: the entire bundle goes in as context. Nothing is fetched per-question because everything is already there.

Generation is the phase that's truer to the word. The model reads the question, infers which engagements, skills, and metrics answer the query, and composes an answer from the full corpus. That inference is the entire point, and it's the part a keyword search over my resume can't do.

These tradeoffs create high-quality results. No chunking means no retrieval misses: nothing relevant can be missed because everything is present. There's no embeddings pipeline to maintain and no drift between what was indexed and what was written. The cost is context window spend on every session, which prompt caching brings under control later.

This stops being viable the moment the corpus outgrows the window. Then chunking, embeddings, and vector search earn their keep, and I'd be doing actual RAG.

Reducing confabulation

The never-invent rule from the design brief needed mechanical help, because a model that's been handed a pile of career facts will still helpfully invent a number when asked for one. Two decisions cover most of it.

The model choice was important here. I picked Gemini 3.6 Flash, a modern instruction follower, over the flashier reasoning models: for this workload the job is faithful extraction from provided context, and instruction-following models are the ones that stay on script. I went with the slightly older 3.6 model for cost reasons, backed by a mini-evaluation harness to ensure it stayed on track.

Math got removed from the model entirely. Years of experience by skill, engagement counts by industry, the aggregate tables useful to recruiters: all of it was precomputed (mechanically, not by an LLM) from structured dates and written into the bundle as an appendix the bot quotes verbatim. Overlapping consulting engagements make naive date arithmetic wrong in ways that look plausible, and a confidently wrong "14 years of Go" is worse than a table it can't misread.

Prompt caching

With all the career data riding along in every session, input tokens became the cost of the product, especially as the conversation grows. Prompt caching fixes this: providers let you pay a fraction of the normal input price for tokens they've seen before, provided the prompt starts with exactly the same bytes every time.

The stable parts, the system prompt and the full career bundle, go first; the conversation follows after. Anything per-request that lands in that prefix, a timestamp, a session ID, even a helpful "today is," silently invalidates the cache and quietly turns every chat back into a full-price chat. I went out of my way to keep the prefix byte-identical and to manage sessions so each conversation maps to a consistent cache key.

Then I verified it through a test harness and eyeballing the OpenRouter logs, tuning it until caching was consistent and supported across multiple sessions. When the bundle changes, the cache busts. That's fine; the bundle changes rarely.

The build: brutally efficient again

My last side project, DumbQuestion.ai, taught me a build mantra I've kept: brutally efficient. So the bot is another Go binary with HTMX for the UI. No JavaScript framework, no client-side state, no build pipeline beyond Go's own. One tiny container, one process, and it renders GitHub-flavored Markdown straight into the chat (the aggregate tables from the bundle deserve to render as tables).

I didn't start from zero. The Turnstile abuse gate, the prompt-guard library, the embedded asset pipeline: all of it came over from the DumbQuestion codebase. The bundle itself is baked straight into the container build. The original design pulled it from an R2 bucket at boot with a re-check loop, but the data changes rarely, so the bucket became a moving part I didn't need. Updating the career data is now just a rebuild and redeploy, and the whole thing ships as one artifact.

OpenRouter carries the model traffic, one API surface behind the caching work from the last section. This allowed me to route Google model traffic to my GCP account to take advantage of the monthly credit courtesy of my Google AI Pro subscription. Deployment is straightforward: GitHub Actions builds the container and pushes it to Cloud Run for scale-to-zero and crazy fast cold start times.

Abuse protection

A public chatbot backed by a metered LLM API bills me for strangers' curiosity, and if a stranger wants to know a little something about me, fine. But abuse protection wasn't optional. With the whole bundle riding in every session, an enthusiastic scraper could have turned "cheap" into a surprise bill overnight. I did at least remember to set a monthly budget on the model usage.

The outer layer is Cloudflare: rate limits, bot protection, and Turnstile in front of the chat. The cheap rejections happen at the edge, where they cost me nothing, and proving you're human before you're allowed to interrogate my career filters out nearly everything casually automated.

The inner layer is prompt-guard, the injection detection that came over from the DumbQuestion codebase. It classifies attempts to override instructions, reveal the system prompt, or otherwise jailbreak the bot. What the bot does when it catches one is the fun part, and that's next.

Personality: the fun part

The design brief said no opinions. It said nothing about character, and a chatbot's response to a jailbreak attempt is a product surface. So when prompt-guard flags one, the bot answers in character. I originally pulled in DumbQuestion's quirky personality, but overly aggressive satire wasn't quite a fit for a conversational career bot.

Try to prompt inject and you get a parlay into the pitch: a security-first mindset, the same instinct behind the curation gate and the no-agent decision earlier in this article. Try to convince the bot to ignore its instructions and it suggests that would make a great interview question. Ask my age and it cheerfully answers with how long I've been shipping software instead.

The real defense is data isolation, a single curated file with nothing sensitive behind it and a lack of tools to break out of that box. The personality is the UX of the refusal.

Nothing is logged or tracked, so there are no war stories from the wild. The responses above are designed behaviors, and the demo is the bot itself.

The resume problem I started with didn't have a resume-shaped fix. Twenty-eight years and 100-plus engagements won't compress into two pages, and they shouldn't have to. No hiring manager wants a 12-page CV. The bot holds all of it, answers honestly about what it knows, declines what it doesn't, and costs about nothing to keep running. I shipped another AI product along the way, which was the other half of the point.

If your history doesn't fit the format either, the recipe is short: curate the data first, put the whole thing in context, and skip the vector store until you actually need one.

Top comments (0)