DEV Community

Shrinidhi Thalla
Shrinidhi Thalla

Posted on

# The Silent Truncation Bug That Almost Broke My Hindsight Agent's Data

Nobody warns you that the hardest part of building an LLM-backed agent isn't the agent
logic — it's everything upstream and downstream of the model call. I learned this the
annoying way while generating test data and querying Hindsight
for a memory-backed sales assistant, and the two bugs that hit me are the kind that don't
show up until you're already relying on the system working.


This is the reliability side of the build that the demo doesn't show.

What I was building

A Deal Intelligence Agent: it stores every sales call as a memory in Hindsight, tagged
by company, then recalls relevant history to brief a rep before their next call or to spot
patterns across every deal at once. The agent logic itself came together fast. What didn't
come together fast was making the two LLM-dependent pieces — data generation and memory
recall — actually reliable.

Bug one: the model was cutting itself off mid-JSON

To test the system, I needed realistic-looking sales data, so I asked an LLM to generate it
directly as JSON — six deals per request, each with several call notes. The first few runs
looked fine. Then, intermittently, I'd get errors like Unterminated string starting at
line 172
or Expecting value: line 126.

The model wasn't malfunctioning — it was running out of output budget before finishing the
JSON. Models like gpt-oss-120b spend part of their response on internal reasoning before
they ever write the answer, and a big structured JSON blob plus reasoning overhead adds up
fast. Ask for too much in one call, and the JSON gets cut off mid-string with no warning
except a broken parse.

The fix was to stop asking for one big thing and start asking for several small things:

def generate(i1, s1, i2, s2):
    prompt = (PROMPT.replace("INDUSTRY1", i1).replace("STAGE1", s1)
                    .replace("INDUSTRY2", i2).replace("STAGE2", s2))
    for _ in range(3):
        try:
            r = groq.chat.completions.create(
                model="openai/gpt-oss-120b",
                messages=[{"role": "user", "content": prompt}],
                reasoning_effort="low",
                max_completion_tokens=6000,
            )
            text = re.sub(r"^```

(?:json)?\s*|\s*

```$", "", r.choices[0].message.content.strip())
            return json.loads(text)["deals"]
        except Exception as e:
            print("  Retrying after error:", e)
    raise SystemExit("Generation failed 3 times")
Enter fullscreen mode Exit fullscreen mode

Three changes did the work: requesting 2 deals per call instead of 6, forcing
reasoning_effort="low" so the model spent less of its budget thinking instead of writing,
and a hard max_completion_tokens ceiling so I knew exactly how much room the reply had.
Smaller, more numerous requests beat one large fragile one.

Bug two: the recall step hit a rate limit mid-demo

The second failure showed up later, live, while testing the cross-deal pattern search:

groq.APIStatusError: Error code: 413 - {'message': 'Request too large for model
`openai/gpt-oss-120b`... Limit 8000, Requested 8004...'}
Enter fullscreen mode Exit fullscreen mode

Groq's free tier caps token throughput per minute, and my recall step was pulling up to
8,000 tokens of memory out of Hindsight, then wrapping it in a prompt on top — enough to
tip over the limit by a handful of tokens. The failure mode was a full crash: a red
traceback, no graceful fallback, right in the middle of what was supposed to be a clean
demo.

The fix had two parts. First, shrink the actual request size so it has headroom instead of
sitting right at the ceiling:

def recall_context(query, max_tokens=3000, budget="mid", tags=None):
    kwargs = dict(bank_id=BANK, query=query, max_tokens=max_tokens, budget=budget)
    if tags:
        kwargs["tags"] = tags
        kwargs["tags_match"] = "all_strict"
    memories = hindsight.recall(**kwargs)
    return getattr(memories, "text", None) or "\n".join(r.text for r in memories.results)
Enter fullscreen mode Exit fullscreen mode

Second, stop letting a limit crash the whole app. If it does happen anyway, show something
a person can read instead of a stack trace:

def ask_llm(prompt):
    try:
        r = groq.chat.completions.create(
            model=MODEL, messages=[{"role": "user", "content": prompt}],
            reasoning_effort="low", max_completion_tokens=1500,
        )
        return r.choices[0].message.content
    except Exception as e:
        return f"⚠️ The language model hit a limit ({type(e).__name__}). Wait a minute and try again."
Enter fullscreen mode Exit fullscreen mode

That one try/except turned a hard crash into a message a user can actually act on — which
matters a lot more once you're demoing something live than it does while you're alone in a
terminal.

What this looked like once fixed

After both fixes, the same pattern-insight query that used to crash now consistently
returns a real answer — naming specific companies from the dataset and drawing a
conclusion, comfortably inside the token budget instead of grazing the ceiling. The
generation script that used to fail roughly half the time on a six-deal batch now completes
cleanly, batch by batch, with retries as a safety net rather than the main plan.

An honest limitation

The retry logic I wrote catches failures and tries again, but it doesn't yet distinguish
why a call failed — a rate limit and a genuine malformed response get the same blind
retry. For a real production system, those need different handling: a rate limit should
back off and wait, not immediately retry into the same wall.

Lessons learned

  • A model running out of output tokens fails silently, not loudly. It doesn't say "I ran out of room" — it just stops, often mid-string, and you get a parser error that looks unrelated to the real cause.
  • Reasoning-capable models spend budget you don't see. reasoning_effort is worth setting explicitly for structured-output tasks instead of trusting the default.
  • Smaller, more frequent requests beat one large one. This is true for both generation and recall — batching down reduced my failure rate more than any amount of prompt tweaking did.
  • A try/except around your LLM call isn't optional once real people click the button. Assume the token limit will be hit eventually, and design for the moment it happens instead of hoping it won't.

If you're building anything on Hindsight or
curious about agent memory in general, the
full project is here: github.com/ghaneeshkumar83-lab/deal-intelligence-agent.

Top comments (0)