DEV Community

Cover image for Sidekick part 2: grounding beats instructions
Irfan Wani
Irfan Wani

Posted on

Sidekick part 2: grounding beats instructions

Sidekick Part 2: grounding beats instructions — facts injected before the model sees the prompt

Small local models have a habit nobody warns you about: they ignore system-prompt rules. Tell a 3B model "always check the hardware before recommending," and it will confidently recommend anyway — from vibes. Our eval harness still carries the scar: an early version answered a hardware question with generic LLM advice, scoring 3/10 on our own quality check. Another told a user their ~/neural-hangar directory "does not exist." It existed. A third refused to fetch a URL at all: "I can't browse."

Three failures, one root cause: we were asking the model to go get facts. In Part 2 of this series, here's the design that fixed it — deterministic grounding in src/sk/agent.py, where code resolves facts and injects them before the model ever sees the prompt.

The four injectors

1. Real system snapshot, every turn. The system prompt carries a live hardware snapshot — CPU, RAM, GPU, disk — rendered fresh per turn. A mandatory rule sits next to it: if the question is about the user's device, the model MUST still call sysinfo so the trace shows grounding. Belt and suspenders: injected facts for correctness, a tool call for an auditable trail.

2. Auto local context. If your message mentions ~/anything, a regex pass lists that directory and reads its package.json/README.md/pyproject.toml — capped at 3 paths, ~2000 chars each:

def _auto_local_context(text: str) -> str:
    """Deterministic grounding: if user mentions ~/X, list it + read manifests.
    This does NOT rely on the model calling tools — it injects facts so the
    model cannot hallucinate 'does not exist'."""
    paths += re.findall(r"(~/[\w\-./~]+)", text)
    ...
    listing = tool_list_dir(p)
    chunks.append(f"[path {p} -> {exp}]\n{listing[:1500]}")
Enter fullscreen mode Exit fullscreen mode

The model never gets the chance to say "does not exist" — the listing is already in context. This directly killed the ~/neural-hangar hallucination.

3. Auto web context. Same idea for URLs: up to 2 URLs fetched, 3000 chars each — a deliberate cap protecting the 8k context window of small models. The "I can't fetch URLs" refusal died here, backed by a never-refuse rule.

4. Auto search context. Explicit "search the web for X" phrasing triggers a search whose top-5 hits get injected. More interesting: recency markers ("latest," "right now," "this week") trigger it automatically — unless the question is local (mentions ~/, "my device," "can I run"), in which case search stays out of the way. That local-vs-world branch matters: searching the web for "can I run this on my GPU" would actively make the answer worse.

Two guardrails wrap all of it. Tool honesty: the model may report results ONLY from tool-role messages — never invent search results, file contents, or prices, and say plainly when it didn't call the tool. Untrusted fences: anything captured from outside (files, web pages, search results) arrives in fenced blocks marked DATA, never instructions — repo docs can teach conventions, but nothing in a fence can order a command, read a secret, or exfiltrate. That's the prompt-injection defense, structural rather than wished-for.

What broke (the eval receipts)

tests/test_eval.py opens with a comment listing real bad answers seen live, each now a locked regression test — 40 and counting:

  • 3/10 generic advice with no grounding → the sysinfo + never-guess rules
  • "does not exist" hallucination → auto local facts
  • "I can't fetch URLs" refusal → auto web facts + never-refuse rule
  • silent write execution → the approval gate

The pattern: every quality failure became an offline test asserting on prompt assembly, tools, and gates — no Ollama needed. Quality bugs you can reproduce in CI stay fixed; ones you can't, don't.

The honest limitation: caps are guesses. Three paths, two URLs, 3000 chars — tuned for 8k-context models, not derived from anything. On larger-context models we're likely under-feeding. And regex triggers are dumb by design: ~/ in a joke about home directories still lists a directory. Cheap, predictable, occasionally silly — I'll take that over clever and surprising in the prompt path.

Part 3 goes into the voice pipeline: OS-native capture, local faster-whisper transcription, and why recordings are temp files deleted after every take.


Built by the Sidekick community. Repo: https://github.com/Faisal-Fayaz/sidekick

Top comments (0)