The alert fires at 3:14 AM. Your LLM service returns 500s. The free quota is gone. You check the dashboard. 10,000,000 tokens. Zero remaining.
Most guides teach prevention. Budgets. Ledgers. Pre-checks. This one teaches survival. What happens after the quota dies? The answer determines if your users stay.
I built a degradation state machine. It turns quota exhaustion into a designed state. Not a crash. This article shows the full implementation.
Disclosure: This article was prepared as part of MonkeyCode's product outreach.
The four states
Free tiers end. That is a fact. Design for it.
| State | Trigger | Behavior |
|---|---|---|
| NORMAL | Usage < 70% | Full service |
| WARNING | Usage >= 70% | Cache responses |
| DEGRADED | Usage >= 90% | Critical requests only |
| EXHAUSTED | Usage >= 100% | Cache or 503 |
The thresholds are config. Tune them for your traffic. The state machine is the core.
Step 1: Track the quota
MonkeyCode's free tier includes 10M tokens and a free server option. Quotas change. Verify the current numbers first.
The tracker polls the usage endpoint. It stores the remaining count in memory.
class QuotaTracker:
def __init__(self, cap: int):
self.cap = cap
self.used = 0
def record(self, tokens: int):
self.used += tokens
def ratio(self) -> float:
return self.used / self.cap
def state(self) -> str:
r = self.ratio()
if r >= 1.0:
return "EXHAUSTED"
if r >= 0.9:
return "DEGRADED"
if r >= 0.7:
return "WARNING"
return "NORMAL"
Simple arithmetic. Four states. No magic.
Step 2: Add the cache
The WARNING state needs a cache. Every response gets stored. Keyed by the prompt hash.
import hashlib
class ResponseCache:
def __init__(self):
self.store = {}
def key(self, prompt: str) -> str:
return hashlib.sha256(prompt.encode()).hexdigest()
def get(self, prompt: str):
return self.store.get(self.key(prompt))
def put(self, prompt: str, response: str):
self.store[self.key(prompt)] = response
The cache is a dictionary. For production, use Redis. For a free server, a dict is fine.
Step 3: Classify requests
Not all requests are equal. Some deserve the last tokens. Others can wait.
def classify_request(payload: dict) -> str:
if payload.get("priority") == "critical":
return "critical"
if payload.get("task") == "classify":
return "cheap"
return "normal"
Critical requests keep the service alive. Cheap requests can use cached answers. Normal requests wait.
Step 4: The router with degradation
This is the heart. The router checks the state before every call.
def route(payload: dict, tracker: QuotaTracker, cache: ResponseCache):
state = tracker.state()
task = classify_request(payload)
if state == "EXHAUSTED":
cached = cache.get(payload["text"])
if cached:
return {"source": "cache", "label": cached}
return {"source": "none", "error": "quota exhausted"}, 503
if state == "DEGRADED" and task != "critical":
cached = cache.get(payload["text"])
if cached:
return {"source": "cache", "label": cached}
return {"source": "none", "error": "degraded"}, 503
if state == "WARNING":
cached = cache.get(payload["text"])
if cached:
return {"source": "cache", "label": cached}
# live call to the model
result = call_model(payload["text"])
cache.put(payload["text"], result)
tracker.record(estimate_tokens(payload["text"], result))
return {"source": "live", "label": result}
Read the logic top to bottom. Each state adds a restriction. The service never crashes. It degrades.
Step 5: Test the degradation
A state machine needs a test. Simulate quota exhaustion. Verify each state.
def test_degradation():
tracker = QuotaTracker(cap=1000)
cache = ResponseCache()
# fill the cache
cache.put("hello", "greeting")
# NORMAL
assert tracker.state() == "NORMAL"
# WARNING
tracker.used = 700
assert tracker.state() == "WARNING"
# DEGRADED
tracker.used = 900
assert tracker.state() == "DEGRADED"
# EXHAUSTED
tracker.used = 1000
assert tracker.state() == "EXHAUSTED"
# EXHAUSTED serves from cache
result = route({"text": "hello"}, tracker, cache)
assert result["source"] == "cache"
# EXHAUSTED rejects uncached
result = route({"text": "unknown"}, tracker, cache)
assert result[1] == 503
print("All degradation tests passed")
Run it.
python test_degradation.py
Step 6: Deploy and monitor
The free server option hosts this. No credit card. No billing alarm.
git clone <your-repo>
cd <your-repo>
pip install -r requirements.txt
uvicorn app:app --host 0.0.0.0 --port 8000
Add a health endpoint that reports the current state.
curl http://localhost:8000/state
# {"state": "NORMAL", "ratio": 0.42}
A simple curl tells you the truth. No dashboard needed.
What the cache cannot save
Caching has limits. Unique prompts miss. Long conversations miss. Time-sensitive answers miss.
The cache hit rate determines your survival time. Measure it.
hits = sum(1 for r in results if r["source"] == "cache")
print(f"Hit rate: {hits / len(results):.0%}")
A 40% hit rate extends the service by days. A 5% hit rate barely helps.
Who should not use this
Teams with SLAs need paid capacity. Teams with real-time requirements need dedicated endpoints. Teams with high concurrency need more than a free server.
This pattern is for prototypes. For internal tools. For services where a 503 is acceptable. For developers who want a service that fails gracefully.
The takeaway
Quota exhaustion is a state. Design for it. The state machine turns a crash into a graceful degradation. Cache what you can. Protect what matters. Fail loudly when you must.
If you want to experiment with this pattern, MonkeyCode's free tier is a place to start. Check the current quota first. Then build something that survives the month.
MonkeyCode provides free models that can run this workflow.
Top comments (0)