DEV Community

jidonglab
jidonglab

Posted on

Ollama num_ctx Truncated 287 of 400 Prompts and Never Told Me

My local RAG bot answered 46% of my test questions correctly. Llama 3.1 8B on Ollama, a 128K context window on the model card, retrieved chunks that I had checked by hand. The right paragraph was in the prompt every single time.

The model just never saw it. Ollama's num_ctx was set to 2048 tokens, and Ollama quietly chopped the front off every prompt longer than that. No error, no field in the response, no warning in my Python logs. 287 of my 400 prompts got cut.

This is the autopsy.

TL;DR

  • Ollama num_ctx is the context window Ollama actually allocates, not the one on the model card. In the version I ran, the default was 2048 tokens even for a 128K model.
  • Overflowing prompts are truncated silently from the front. The tail survives, so the question stays and the instructions and top-ranked chunks vanish.
  • Spot it with prompt_eval_count in the response. If it sits at or just under your num_ctx no matter how long your prompt is, you're being truncated.
  • Fix it with options: {"num_ctx": 8192} on the native API, or PARAMETER num_ctx 8192 in a Modelfile. The OpenAI-compatible endpoint ignored my per-request setting.
  • After the fix: accuracy went from 46% to 81% on the same 400 questions. Cost: about 1 GB more VRAM and a slower median response.

What was I building?

A notes assistant. Around 3,000 markdown files from five years of work logs, chunked at roughly 400 tokens, embedded, top 10 chunks retrieved per question. Everything runs locally on a 12 GB GPU because I don't want my work notes leaving the machine.

The prompt went to /api/generate as one string, in this order:

  1. Instructions ("answer only from the context, reply in JSON with answer and source")
  2. The 10 retrieved chunks, best match first
  3. The question

That ordering is the textbook layout. It's also exactly the wrong layout for what Ollama was doing to me.

Why did the model ignore my instructions?

It ignored them because it never received them. When a prompt exceeds num_ctx, Ollama keeps the end of the input and discards tokens from the beginning to make it fit. My instructions lived at the beginning. So did chunk #1, the best match.

The symptoms looked like a dumb model, not a broken pipeline:

  • Answers came back as plain prose instead of JSON, about a third of the time.
  • The source field, when present, pointed at low-ranked chunks.
  • Questions with obvious answers got confident, wrong replies built from chunks 7 through 10.

I spent an evening rewriting the system prompt. I added "IMPORTANT" in caps. I moved the JSON instruction to two places. Nothing changed, which in hindsight was the clue: the edits were landing in a region the model never read.

How do you detect Ollama prompt truncation?

Check prompt_eval_count in the response. It's the number of prompt tokens the model actually processed. If it's pinned at your context limit while your input keeps growing, Ollama is truncating.

I logged it next to my own token estimate for every request:

import requests

def ask(prompt: str, num_ctx: int | None = None) -> dict:
    body = {"model": "llama3.1:8b", "prompt": prompt, "stream": False}
    if num_ctx:
        body["options"] = {"num_ctx": num_ctx}
    r = requests.post("http://localhost:11434/api/generate", json=body, timeout=120)
    r.raise_for_status()
    return r.json()

resp = ask(prompt)
print(len(prompt) // 4, resp["prompt_eval_count"])
Enter fullscreen mode Exit fullscreen mode

The output was embarrassing in its regularity:

5120 2047
3890 2047
1410 1402
6230 2047
Enter fullscreen mode Exit fullscreen mode

Every long prompt was capped at the same number. The short ones passed through intact. The server log had the confession all along, a line containing truncating input prompt with the limit and the original length. I just never read the server log, because the API returned HTTP 200 with a perfectly normal-looking body.

How many of my prompts were affected?

287 of 400, or 72%. With 10 chunks of about 400 tokens plus instructions, most prompts landed between 4,000 and 6,000 tokens. Only the questions whose retrieval returned short chunks fit under 2048.

Splitting the eval by that line made the problem obvious:

Prompt length Questions Correct
Under 2048 tokens 113 96 (85%)
Over 2048 tokens 287 88 (31%)
Total 400 184 (46%)

The model wasn't bad at my task. The model was great at my task whenever it could see the task.

How do you set num_ctx in Ollama?

Pass it per request in options on the native API, or bake it into a model with a Modelfile. Both worked for me. One thing that did not work is the obvious one.

Option 1: per request on the native API

resp = ask(prompt, num_ctx=8192)
Enter fullscreen mode Exit fullscreen mode

Option 2: a Modelfile

FROM llama3.1:8b
PARAMETER num_ctx 8192
Enter fullscreen mode Exit fullscreen mode
ollama create notes-llama -f Modelfile
Enter fullscreen mode Exit fullscreen mode

Then call notes-llama instead of llama3.1:8b. This is the version I kept, because it makes the setting impossible to forget in some other script.

What failed: the OpenAI-compatible endpoint

My first fix was through /v1/chat/completions with the OpenAI Python client, sending num_ctx via extra_body. prompt_eval_count (exposed as usage.prompt_tokens) stayed at 2047. That endpoint didn't honor the option in my setup. If you use Ollama behind an OpenAI-style client, set the context in a Modelfile or at the server level. Newer Ollama releases added an OLLAMA_CONTEXT_LENGTH environment variable for the server-wide default, and raised the default itself, so check what your version does instead of trusting either my number or the model card.

What did the bigger context window cost?

About 1 GB of VRAM and roughly 1.5 seconds of median latency. The KV cache scales linearly with num_ctx, so a 4x window costs 4x the cache.

For Llama 3.1 8B at fp16 the math is short: 32 layers × 8 KV heads × 128 dims × 2 (K and V) × 2 bytes = 128 KB per token.

  • 2048 tokens: 256 MB of KV cache
  • 8192 tokens: 1 GB of KV cache

ollama ps moved from about 5.9 GB to 6.8 GB on my machine, which still fits the 12 GB card with room for the embedding model. Median end-to-end time per question went from 1.9 s to 3.4 s, because the model now actually evaluates 5,000 prompt tokens instead of 2,000. That's not a regression. That's the price of reading the whole prompt, which I thought I was already paying.

What were the results after the fix?

81% correct, 324 of 400, same questions, same retrieval, same prompt text. JSON format compliance went from about two thirds to every response but 3. The source field started pointing at chunk #1 or #2 most of the time, which is what you'd hope for when your retriever is any good.

Not everything was fixed. The remaining 76 misses were real retrieval failures and a few questions my notes genuinely can't answer. Those are honest problems. The 140 extra correct answers were never a model problem.

What should you change in your own Ollama setup?

Three habits, all cheap:

  1. Assert on prompt_eval_count. If it's within a few tokens of num_ctx, raise or at least log loudly. One if statement would have saved me an evening of prompt engineering against a wall.
  2. Put the question and the critical instructions at the end of single-string prompts. If something gets cut, it should be the lowest-ranked chunk, not your output format.
  3. Set num_ctx explicitly everywhere. Model card context length is a ceiling, not a default. Pick a number your VRAM can carry and write it down in a Modelfile.
def ask_checked(prompt: str, num_ctx: int = 8192) -> dict:
    resp = ask(prompt, num_ctx=num_ctx)
    if resp["prompt_eval_count"] >= num_ctx - 8:
        raise RuntimeError(f"prompt truncated at {resp['prompt_eval_count']} tokens")
    return resp
Enter fullscreen mode Exit fullscreen mode

So why did Ollama truncate my prompts without telling me?

Ollama truncated 287 of my 400 prompts because num_ctx, the context window it actually allocates, defaulted to 2048 tokens in my version, regardless of the model's advertised 128K. When a prompt exceeds num_ctx, Ollama drops tokens from the front, keeps the tail, and returns a normal HTTP 200; the only traces are a server log line and a prompt_eval_count stuck near the limit. Setting num_ctx to 8192 via request options or a Modelfile PARAMETER (not the OpenAI-compatible endpoint, which ignored it for me) raised my RAG accuracy from 46% to 81% for about 1 GB of extra VRAM.


Written by the developer behind Preterview, an interview prep platform.

Top comments (0)