I built BOOTH, a small checkpoint layer that sits between your app and an
LLM call and returns a structured decision instead of just fluent text.
Example:
Evidence: "Returns are allowed within 45 days."
LLM answer: "Returns are allowed within 90 days."
Confident, fluent, and unsupported. Self-reported confidence won't catch it.
Checking against the evidence will.
BOOTH doesn't decide what's true. You supply the comparison, and BOOTH turns
the outcome into an explicit status you can't accidentally ignore:
result = booth.check_with_evidence(
answer=llm_answer,
evidence=retrieved_docs,
compare_fn=your_comparison_function, # your logic, your definition of "agrees"
)
if result.status == booth.BLOCKED:
print(result.detail) # why it was blocked
Results also carry a reason and a checker_failed flag, so "the answer disagreed"
and "my comparison function crashed" are never mixed up.
There's also check() / acheck() for plain LLM calls: ambiguity detection plus
confidence checks with reconsideration retries.
It keeps your RAG pipeline and provider as they are. Zero runtime dependencies,
provider-agnostic, MIT, Python 3.9+.
pip install boothpy
GitHub: https://github.com/Vedantgitbot/booth
How are you currently deciding whether an LLM output is safe to pass downstream?
Beta · 300+ tests · CI passing · MIT · Python 3.9+
Top comments (0)