A bug that crashes gets noticed right away. A bug that returns a plausible but wrong result can run in production for months without anyone noticing.
While auditing the code of OGX, an open-source framework for building AI agents, I found three of them. None raises an error. Each one quietly skews a result: search relevance, a document filter, or a token bill.
Context: OGX, formerly Llama Stack
OGX is the continuation of Meta's Llama Stack project. It provides the standard building blocks of an agentic application: inference, an OpenAI-compatible Responses API, and vector stores for RAG (Elasticsearch, pgvector, Qdrant…).
Teams that build agents on this kind of foundation trust those blocks without rereading them. That is exactly where an audit pays off: looking for what "almost" works.
Bug 1: one extra letter disables a search setting
Elasticsearch hybrid search merges a keyword search and a vector search (RRF). The rank_window_size parameter sets how many candidates enter that merge.
In the code, the list of allowed parameters contained rank_windows_size, with an extra "s". The correct parameter was treated as unknown and silently dropped.
# before
allowed_rrf_params = {"rank_constant", "rank_windows_size", "filter"}
# after
allowed_rrf_params = {"rank_constant", "rank_window_size", "filter"}
| Requested | Actually sent (before) | After fix |
|---|---|---|
rank_window_size = 50 |
10 (default) | 50 |
A one-character fix, submitted in pull request #6697.
Bug 2: pgvector filters that never match booleans
With pgvector you can filter documents on their metadata, for example "published = true". For "in list" (IN) and "not in list" (NOT IN) filters, the code converted each value to Python text: True became "True".
PostgreSQL extracts the JSON metadata as "true". The two never match. Same problem for a number stored as 2023.0 filtered on 2023.
Tested on a real PostgreSQL 16 database with two documents:
| Filter | Before | After fix |
|---|---|---|
| published in [true] | no result | document 0 |
| published not in [true] | documents 0 and 1 | document 1 |
| year in [2023] | no result | document 1 |
The worst case is the exclusion filter: it lets through documents you meant to exclude. In a RAG pipeline, that means unpublished or outdated sources injected into the model's context. The fix applies to lists the same typed casts already used for single-value comparisons.
Fix submitted in pull request #6698.
Bug 3: a token counter that forgets earlier calls
An agent often makes several model calls for a single response, for example when it uses tools. The Responses API correctly sums input and output tokens across those calls.
But two detailed counters were overwritten instead of summed: cached tokens and reasoning tokens. Only the last call counted.
| Example over 2 calls | Before | After fix |
|---|---|---|
| Cached tokens (80 + 120) | 120 | 200 |
| Reasoning tokens (10 + 15) | 15 | 25 |
These are exactly the two figures used for cost tracking. Caching lowers the bill, reasoning raises it. A dashboard built on these values under-estimated both.
Reported in issue #6699, with a proposed fix (commit e8177d0).
What this means for teams shipping agents
- Test what is silent. An ignored parameter or a filter that doesn't filter triggers no alert. You need a test that checks the value actually sent, not just the absence of errors.
- Test with real types. Bugs 2 and 3 only show up with booleans, floats or chained calls. "Strings and a single call" test suites miss them.
- Don't blindly trust the foundation. A popular open-source framework is still code written fast. Reviewing the critical blocks (filtering, costs) before production takes a few hours.
- Contribute back. Each fix here is a few lines. Submitting it upstream helps everyone who builds on the same foundation.
Building agents on OGX or another framework? I'd love to hear which silent bugs you've run into.
Let's talk, or get your own agents audited: LinkedIn · Malt · Medium
Top comments (0)