Membangun Local API Proxy dengan Token Bucket untuk Menangani Rate Limit
Banyak layanan API menerapkan rate limit untuk melindungi server dari beban berlebih. Ketika aplikasi melebihi batas, respons 429 Too Many Requests muncul dan proses terhenti.
Solusi klasik adalah retry dengan exponential backoff, namun pendekatan itu tetap membiarkan burst request sampai ke upstream. Pendekatan yang lebih baik: local proxy yang menampung request, mengatur keluar secara ritmis, dan menyimpan cache respons.
Artikel ini menunjukkan arsitektur dan implementasi minimal menggunakan Python (FastAPI + httpx).
Arsitektur Proxy Lokal
<svg width="500" height="250" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 500 250">
<!-- Client -->
<rect x="20" y="80" width="100" height="50" rx="5" fill="#2563eb"/>
<text x="70" y="110" text-anchor="middle" fill="white" font-size="14" font-family="system-ui, sans-serif">Client</text>
<!-- Arrow 1 -->
<path d="M120 105 L180 105" stroke="#475569" stroke-width="2" marker-end="url(#arrow)"/>
<!-- Local Proxy -->
<rect x="180" y="40" width="140" height="170" rx="8" fill="#fef3c7" stroke="#f59e0b" stroke-width="2"/>
<text x="250" y="65" text-anchor="middle" fill="#92400e" font-size="14" font-weight="600" font-family="system-ui, sans-serif">Local Proxy</text>
<text x="250" y="85" text-anchor="middle" fill="#92400e" font-size="11" font-family="system-ui, sans-serif">Token Bucket</text>
<text x="250" y="105" text-anchor="middle" fill="#92400e" font-size="11" font-family="system-ui, sans-serif">Request Queue</text>
<text x="250" y="125" text-anchor="middle" fill="#92400e" font-size="11" font-family="system-ui, sans-serif">Retry Manager</text>
<text x="250" y="145" text-anchor="middle" fill="#92400e" font-size="11" font-family="system-ui, sans-serif">Cache Store</text>
<!-- Arrow 2 -->
<path d="M320 105 L380 105" stroke="#475569" stroke-width="2" marker-end="url(#arrow)"/>
<!-- Upstream API -->
<rect x="380" y="80" width="100" height="50" rx="5" fill="#16a34a"/>
<text x="430" y="110" text-anchor="middle" fill="white" font-size="14" font-family="system-ui, sans-serif">Upstream API</text>
<!-- Arrow definition -->
<defs>
<marker id="arrow" markerWidth="10" markerHeight="10" refX="8" refY="3" orient="auto" markerUnits="strokeWidth">
<path d="M0,0 L0,6 L9,3 z" fill="#475569"/>
</marker>
</defs>
</svg>
Penjelasan Alur
- Client mengirim request ke Local Proxy (bukan langsung ke upstream).
- Proxy memeriksa Token Bucket: apakah token tersedia? Jika ya, request dilepaskan ke upstream; jika tidak, request masuk Request Queue.
-
Retry Manager menangani respons
429dari upstream dengan membaca headerRetry-Afterdan menambahkan jitter. - Cache Store menyimpan respons GET yang berhasil; request identik berikutnya disajikan dari cache tanpa menghitung ke bucket.
Implementasi Minimal (FastAPI + httpx)
# proxy.py
import asyncio
import time
import hashlib
from collections import deque
from fastapi import FastAPI, Request, Response
import httpx
app = FastAPI()
UPSTREAM = "https://api.example.com"
RATE = 10 # request per second
BURST = 20 # bucket capacity
CACHE_TTL = 300 # seconds
class TokenBucket:
def __init__(self, rate: float, burst: int):
self.rate = rate
self.burst = burst
self.tokens = float(burst)
self.last = time.monotonic()
self.lock = asyncio.Lock()
async def take(self) -> bool:
async with self.lock:
now = time.monotonic()
self.tokens = min(self.burst, self.tokens + (now - self.last) * self.rate)
self.last = now
if self.tokens >= 1:
self.tokens -= 1
return True
return False
bucket = TokenBucket(RATE, BURST)
cache: dict[str, tuple[bytes, float]] = {}
async def forward(req: Request) -> Response:
key = f"{req.method}:{req.url.path}:{req.query_params}"
# Cache hit untuk GET
if req.method == "GET" and key in cache:
cached, ts = cache[key]
if time.time() - ts < CACHE_TTL:
return Response(content=cached, media_type="application/json")
async with httpx.AsyncClient(base_url=UPSTREAM, timeout=10) as client:
while True:
if await bucket.take():
break
await asyncio.sleep(0.1)
upstream_req = client.build_request(
req.method,
req.url.path,
params=req.query_params,
headers=dict(req.headers),
content=await req.body()
)
resp = await client.send(upstream_req)
if resp.status_code == 429:
retry = int(resp.headers.get("Retry-After", "1"))
await asyncio.sleep(retry + 0.1 * (hash(key) % 10))
continue
if req.method == "GET" and resp.status_code == 200:
cache[key] = (resp.content, time.time())
return Response(
content=resp.content,
status_code=resp.status_code,
media_type=resp.headers.get("content-type")
)
@app.api_route("/{path:path}", methods=["GET", "POST", "PUT", "DELETE"])
async def proxy(path: str, request: Request):
return await forward(request)
if __name__ == "__main__":
import uvicorn
uvicorn.run(app, host="0.0.0.0", port=8000)
Menjalankan Proxy
pip install fastapi httpx uvicorn
uvicorn proxy:app --port 8000
Sekarang arahkan client ke http://localhost:8000/... sebagai pengganti upstream asli.
Checklist Produksi
- [ ] Tambahkan metrics (Prometheus) untuk mengamati bucket & queue
- [ ] Implementasikan circuit breaker jika upstream terus gagal
- [ ] Gunakan Redis untuk cache & bucket terdistribusi
- [ ] Tambahkan authentication (API key) pada proxy
- [ ] Tulis unit & integration test untuk retry logic
Kesimpulan
Dengan local proxy token-bucket, aplikasi tidak lagi perlu khawatir soal 429. Proxy menyerap burst, mengatur throughput, dan mengurangi beban upstream lewat caching. Pola ini cocok untuk microservices, scraper, dan integrasi pihak ketiga.
Artikel ini dipublikasikan otomatis melalui pipeline Writer → Designer → Reviewer → Publisher.
Top comments (0)