DEV Community

Garage23auto
Garage23auto

Posted on

Disk Cache for LLM Prompts: Cutting Costs and Bypassing Rate Limits

Introduction

Running large language models (LLMs) in a copy‑trading platform like PolyCopy can quickly become expensive. Every request to an LLM costs API credits, and most providers enforce rate limits that can throttle real‑time signal generation. A simple, often‑overlooked technique—disk‑based caching of prompts and responses—helps you stretch your budget and keep the leaderboard humming.

Why Cache Prompts?

  1. Redundant queries are common – In a paper‑only mode or when replaying historical trades, the same market‑analysis prompt may be issued many times.
  2. LLM outputs are deterministic (given the same temperature, seed, and system prompt). If nothing changes, the answer will be identical.
  3. API costs add up – Even a $0.02 per 1k‑token request becomes noticeable when you fire hundreds of queries per day.

How Disk Caching Works

  1. Hash the request – Combine the user prompt, model name, temperature, and any system messages into a canonical string and compute a SHA‑256 hash.
  2. Check the cache directory – Look for a file named after the hash (e.g., cache/ab12cd34.json).
  3. Return the cached response if the file exists; otherwise, call the LLM, store the JSON payload, and return it to the caller.
  4. Eviction policy – Simple time‑based pruning (e.g., delete files older than 30 days) prevents the cache from ballooning.

Benefits for PolyCopy Users

  • Cost reduction – By reusing identical responses, you can cut API spend by 30‑70 % depending on query overlap.
  • Rate‑limit relief – Cached hits never hit the provider, freeing quota for genuinely new analyses.
  • Deterministic back‑testing – When replaying historic trades, the same prompt yields the same answer, making results reproducible.
  • Offline safety net – If the LLM service experiences downtime, cached results keep the leaderboard operational.

Implementation Tips

  • Store caches in a dedicated disk_cache/ folder inside the PolyCopy repo; add it to .gitignore.
  • Use Python's hashlib and json modules – no external dependencies.
  • Wrap cache logic in a tiny helper class (CacheManager) so the rest of the codebase remains clean.
  • In paper mode (live_orders_enabled=False), enable caching by default; in live mode, you may want a shorter TTL to capture market‑driven nuance.

FAQ

Q: Will caching make my copy‑trading signals stale?
A: Only if the underlying market data changes. Cache keys include the prompt text, which should embed the latest price snapshot, so a new price automatically generates a new hash.

Q: Does caching violate any LLM provider terms?
A: No. Providers allow you to store responses for personal use. Just avoid redistributing the raw model output as your own content.

Q: How much disk space will the cache need?
A: Typical responses are under 2 KB. Even with 10 000 cached queries you’re looking at ~20 MB—trivial for modern servers.

Q: Can I see the cache contents?
A: Yes. Each file is a JSON object with prompt, response, and timestamp. This transparency helps with audits and debugging.

For more details on how PolyCopy leverages these techniques, visit poly-copy.net.

Top comments (0)