Token-Budget Dry-Run CLI: Preemptively Detecting Context Overflows in Local LLMs
1. Why a CLI (Dry-Run) Instead of a Resident Server?
Setting up a bloated web server merely to validate the token count of an LLM prompt is objectively nonsensical. What we actually need is a lightweight mechanism integrated directly into our CI workflows or local Git hooks—something capable of instantly validating our token budget before a commit or build is ever triggered.
The requirements for this tooling were strictly defined:
- Instantly execute the target model's tokenizer locally against a specified set of input files and a system prompt.
- Calculate and visualize the actual token count and remaining budget on a per-file basis, resolving in milliseconds.
- Guarantee a structured JSON response and an appropriate exit code (Exit Code 0 or 1) in all exceptional cases, including missing dependencies or invalid arguments.
2. Practical Pitfalls and Technical Solutions During Development
When integrating this tool into a CI/CD pipeline, the first major hurdle I encountered was a pragmatic conflict: the default behavior of CLI frameworks versus the necessity of strict JSON outputs.
The Trap: The Default Error Handling of argparse
Python's standard argparse library is undeniably powerful, but it harbors a fatal flaw for automated pipelines. When a type conversion error occurs (e.g., passing a string to --max-tokens) or a required argument is omitted, argparse automatically dumps a plain-text error message to standard error (stderr) and abruptly terminates the process via sys.exit(2).
When invoking this CLI from an automated pipeline or a wrapper script, receiving unstructured plain text upon every parsing failure makes downstream error handling virtually impossible. I needed to enforce a strict contract: regardless of the error type, the tool must respond with machine-readable JSON and terminate with predictable exit codes.
The Solution: Explicit Exception Catching and Enforced JSON Output
To circumvent this default behavior, I deliberately altered the design to bypass strict type conversion during the initial argparse parsing phase. Instead, arguments are received as raw strings and subjected to manual validation. Furthermore, all anomalous flows—including the failure to import critical dependencies like tiktoken—are wrapped in comprehensive try-except blocks. This ensures that every failure routes through a unified JSON output function before terminating the process.
3. The Production-Ready Code
Below is the finalized script, refined to handle numerous edge cases and elevated to a standard capable of withstanding the rigors of production CI/CD environments.
import sys
import json
import argparse
from pathlib import Path
# Explicitly catch missing dependencies and terminate with a JSON error
try:
import tiktoken
except (ImportError, ModuleNotFoundError) as e:
print(json.dumps({"error": "Missing dependency", "message": f"Failed to import 'tiktoken': {str(e)}"}))
sys.exit(1)
def exit_with_json(error_type, message, code=1):
"""Unified error termination handler for pipeline integration."""
print(json.dumps({"error": error_type, "message": message}))
sys.exit(code)
def check_tokens(file_paths, system_prompt, max_tokens):
encoder = tiktoken.get_encoding("cl100k_base")
system_tokens = len(encoder.encode(system_prompt))
results = {"status": "PASS", "details": []}
total_fail = False
for file_path in file_paths:
path = Path(file_path)
if not path.exists():
results["details"].append({"file": str(path), "error": "File not found", "status": "ERROR"})
total_fail = True
continue
try:
content = path.read_text(encoding='utf-8')
content_tokens = len(encoder.encode(content))
current_total = system_tokens + content_tokens
budget_remaining = max_tokens - current_total
is_over = budget_remaining < 0
if is_over:
total_fail = True
results["details"].append({
"file": str(path),
"token_count": content_tokens,
"total_with_system": current_total,
"remaining": budget_remaining,
"status": "FAIL" if is_over else "OK"
})
except Exception as e:
results["details"].append({"file": str(path), "error": str(e), "status": "ERROR"})
total_fail = True
results["status"] = "FAIL" if total_fail else "PASS"
return results, total_fail
def main():
parser = argparse.ArgumentParser(description="Token-Budget Dry-Run CLI")
parser.add_argument("--files", nargs='+', required=True)
parser.add_argument("--system-prompt", required=True)
parser.add_argument("--max-tokens", default="8192")
args = parser.parse_args()
# Bypass the default argparse behavior (sys.exit(2)) to guarantee JSON output
try:
max_tokens = int(args.max_tokens)
except ValueError:
exit_with_json("Invalid Arguments", "--max-tokens must be an integer")
results, is_fail = check_tokens(args.files, args.system_prompt, max_tokens)
print(json.dumps(results, indent=2))
sys.exit(1 if is_fail else 0)
if __name__ == "__main__":
main()
💡 For immediate deployment: The complete source code suite (ZIP) for this architecture is available on Gumroad for $0+ (Pay What You Want).
4. Execution Example and CI Integration
Let's examine the behavior of this script when executed in a terminal environment.
$ python token_budget_checker.py \
--system-prompt "You are a senior code reviewer." \
--max-tokens 4096 \
--files src/main.py src/utils.py
If the token counts remain within the specified budget, the script outputs structured JSON to standard output (stdout) and returns an Exit Code of 0.
{
"status": "PASS",
"details": [
{
"file": "src/main.py",
"token_count": 1250,
"total_with_system": 1257,
"remaining": 2839,
"status": "OK"
},
{
"file": "src/utils.py",
"token_count": 820,
"total_with_system": 827,
"remaining": 3269,
"status": "OK"
}
]
}
Conversely, if the token count exceeds the upper limit for any given file, its status flips to "FAIL", and the overarching process terminates with an exit code of 1. This deterministic behavior allows CI pipelines (such as GitHub Actions) to immediately halt and fail the corresponding build step.
Conclusion
In local LLM development, token management often degenerates into an opaque "black box" where limits are only discovered at runtime. However, much like managing traditional infrastructure, the bloat of source code and prompts can be preemptively detected and governed through pre-build static analysis (dry-runs).
While the script presented here is fundamentally simple, integrating it into your operational pipeline establishes a robust defense against the silent truncation risks inherent to LLM workflows. By substituting the encoding logic to match your target model (e.g., Llama, Mistral), you can effortlessly customize this architecture to align with your specific CI/CD constraints.
If this engineering log saved your production server (and your sanity), consider supporting our architecture on GitHub Sponsors.
Top comments (0)