At Idyllic Services, I was working on an automated candidate sourcing system running LLM calls at scale. One iteration consumed 200,000+ tokens (costing ₹15–20 per run) and took 6 minutes.
I redesigned the pipeline architecture to cut token consumption by ~99% down to 5,000 tokens and reduced execution time to 90 seconds. Here is the exact 3-step breakdown of how I did it.
1. RAG Integration, Semantic Routing & Caching
The biggest inefficiency was dumping full candidate histories and long job descriptions into every single prompt.
- RAG Memory Retrieval: I integrated Qdrant vector database to store conversation history and candidate data. Instead of sending 15,000 tokens of raw history, nearest-neighbor search retrieved only the top 2–3 relevant snippets (~400 tokens) per step.
- Semantic Routing: Instead of sending every request to expensive LLMs, I implemented semantic routing to direct deterministic tasks to fast regex or local Python scripts.
- SHA-256 Response Caching: For repeat queries (same JD, same candidate profile), I cached responses using prompt content hashes — making duplicate runs cost 0 tokens.
2. TOON Format (Token-Oriented Object Notation)
JSON has massive syntactic overhead (braces, quotes, repeated keys). When sending structured arrays of candidate data to LLMs, this overhead is pure waste.
I switched the prompt payload format to TOON (Token-Oriented Object Notation), which declares keys once as a header row:
# JSON (before) — 847 tokens for 5 candidates:
[
{"name": "Rahul Sharma", "skills": ["Python", "AWS"], "yoe": 4},
{"name": "Priya Patil", "skills": ["React", "Node"], "yoe": 2}
]
# TOON (after) — 290 tokens for 5 candidates:
{name | skills | yoe}
Rahul Sharma | Python, AWS | 4
Priya Patil | React, Node | 2
Result: 30–40% baseline token reduction on all structured payloads with zero information loss.
3. Removing LLMs from Loops · Batch API Calls
The original codebase called the LLM once per candidate inside a for loop (50 candidates = 50 separate API calls, each carrying full instructions).
I replaced the sequential loop with a single structured batch prompt:
# BEFORE: LLM inside a loop (50 API calls = 200K tokens)
for candidate in candidates:
score = llm.invoke(f"Score candidate:\nJD: {full_jd}\nCandidate: {candidate}")
# AFTER: Single batched call (1 API call = ~5K tokens)
batch_prompt = build_batch_prompt(jd_toon, candidates_toon)
all_scores = llm.invoke(batch_prompt)
Summary of Results
| Optimization | Impact |
|---|---|
| RAG + Routing + Caching | -95% memory tokens, 0 tokens on repeats |
| TOON Format | -30% to -40% payload overhead |
| Batch API Calls | 50 calls → 1 call (Runtime: 6 min → 90 sec) |
Final Outcome: 200,000 tokens → 5,000 tokens (~99% reduction).
Written by Rugved Chandekar (rugved.me)
Top comments (2)
The "LLM inside a for loop" one is the fix that pays back the fastest in my experience — collapsing 50 per-candidate calls into one structured batch prompt usually beats both the RAG and format wins combined, because you stop re-sending the instruction block 50 times. The TOON swap is clever; I've gotten a similar effect just by flattening structured arrays to CSV-ish rows with a header, and the models parse it fine.
One thing I'd watch on the SHA-256 cache: it's a clean win for truly identical (JD, candidate) pairs, but the moment any part of the prompt carries a timestamp, run id, or reordered list, your hit rate quietly collapses and you're paying full price while thinking you're cached. We ended up hashing a normalized/canonicalized payload rather than the raw string. Did you hit any cache-key drift like that, or were your inputs stable enough that raw hashing held up?
Some comments may only be visible to logged-in visitors. Sign in to view all comments.