DEV Community

AGIRouter
AGIRouter

Posted on

How I cut LLM batch costs with time-of-use scheduling

Every batch job I run against LLM APIs used to cost the same at 2 a.m. as at 2 p.m. Then I started routing workloads through a unified gateway that charges different rates at different hours, and my overnight evaluation runs dropped to half the token cost — without touching a single prompt. This post shows the scheduling pattern, the math, and the pitfalls, so you can decide whether it fits your stack.

Disclosure: I'm a founder of AGIRouter, the gateway behind the numbers below. Every pricing figure is an example taken from the platform's public pricing API (agirouter.org/api/pricing) and may change — check the official pricing page before relying on any number. This article was drafted with AI assistance and reviewed and fact-checked by me before publishing.

What time-of-use pricing actually means here

AGIRouter is a multi-model AI gateway built on the open-source New-API project: one API key reaches 29 models across the DeepSeek, Kimi, GLM, Ant Group Ling, GPT-5.5/5.6/6, and Gemini-3 families through a single OpenAI-compatible endpoint (it also speaks the Anthropic, Gemini, and image-generation protocols), imports into Cherry Studio or Lobe Chat with one click, and bills in USD via Stripe.

Some models on the platform use tiered, time-of-use billing (the billing_expr field in the pricing API). For the DeepSeek v4 series, the day splits into two tiers in Asia/Shanghai time:

  • Off-peak (谷价): 00:00–01:00, 04:00–06:00, and 10:00–24:00 — 17 hours a day
  • Peak (峰价): 01:00–04:00 and 06:00–10:00 — 7 hours a day

The off-peak tier is exactly 50% of the peak tier for the DeepSeek models I checked. To be clear: this is AGIRouter's platform pricing strategy, not an official price cut from the model vendors.

Why batch jobs are the perfect candidate

Interactive traffic needs the answer now, so it pays peak rates. Batch work — evaluation sets, embedding generation, migration scripts, bulk summarization, dataset labeling — is delay-tolerant. If a job can wait a few hours, it can wait for the off-peak window. That's the entire idea: separate when a job runs from what the job does.

The scheduler (Python)

import time
from datetime import datetime
from zoneinfo import ZoneInfo

SHANGHAI = ZoneInfo("Asia/Shanghai")

def in_off_peak(now=None):
    """AGIRouter off-peak (谷价) windows, Asia/Shanghai: 00-01, 04-06, 10-24."""
    h = (now or datetime.now(SHANGHAI)).hour
    return h == 0 or 4 <= h < 6 or h >= 10

def seconds_until_off_peak(now=None):
    """Seconds until the next off-peak window opens (0 if already off-peak)."""
    now = now or datetime.now(SHANGHAI)
    h = now.hour
    if h == 0 or 4 <= h < 6 or h >= 10:
        return 0
    target = 4 if h < 4 else 10
    nxt = now.replace(hour=target, minute=0, second=0, microsecond=0)
    return (nxt - now).total_seconds()

def run_batch(prompts, client, model="deepseek-v4-flash"):
    wait = seconds_until_off_peak()
    if wait:
        print(f"Peak window: sleeping {wait/3600:.1f}h before {len(prompts)} job(s)")
        time.sleep(wait)
    return [client.chat.completions.create(
        model=model,
        messages=[{"role": "user", "content": p}],
    ) for p in prompts]
Enter fullscreen mode Exit fullscreen mode

Two production notes. First, time.sleep is fine for scripts, but for real queues use Celery beat, cron, or a cloud scheduler with retries and idempotency keys. Second, test the timezone logic hard: Asia/Shanghai is UTC+8 with no DST, but your servers probably run UTC, and an off-by-eight-hours bug silently converts savings into peak-rate bills.

What the numbers look like

The table below uses the pricing multipliers published in the API for deepseek-v4-flash (p = input token price, c = output token price, cr = cache-hit price). These are relative price units, not currency — the platform converts them into actual charges — so treat the table as an example and check the pricing page for current figures.

Workload (1M input + 500k output tokens) Peak Off-peak
Fresh input only 900,000 450,000
80% of input served from cache 664,800 332,400

Same job, same model, same output quality — the off-peak run costs half. The pro tier behaves identically in relative terms: deepseek-v4-pro is priced at p×1.32 / c×3.96 at peak and p×0.66 / c×1.98 off-peak, so the same 1M+500k run falls from 3,300,000 to 1,650,000 units.

Cache hits: the quiet multiplier

The pricing API also prices cache hits (cr) at roughly 2% of fresh input — 0.003 versus 0.15 units off-peak for v4-flash. Stack the two effects and a cached token processed off-peak costs about 1% of a fresh token processed at peak (0.003 versus 0.30 units). In practice, raising your cache-hit rate means keeping system prompts and document prefixes stable across a batch, grouping similar documents together, and not rewriting prompts mid-run. How much caching a model actually grants depends on the provider and model — the pricing page lists a cache ratio per model where applicable.

Limitations worth knowing

  • Not every model is tiered. In the current pricing API, GLM and Ling models use flat per-tier pricing; only some models (like the DeepSeek v4 series) split by hour. Check each model's billing mode before planning around off-peak windows.
  • Peak hours still matter. Latency-sensitive traffic can't wait; the savings apply only to the slice of work that can.
  • A sleeping job is a delayed job. Add monitoring, retries, and alerting — a batch that silently never runs is worse than an expensive one.
  • Prices and rosters change. The model list and multipliers are versioned in the API (pricing_version); re-verify every billing cycle.
  • Free tiers exist too. Ling-3.1-flash is currently marked free on the platform, which can beat both tiers for suitable workloads.

Bottom line

If your LLM spend includes delay-tolerant batch work, the optimization is low-effort: keep your code identical, move execution into the off-peak window, and keep prompts cache-friendly. At the rates currently published, that pattern cuts DeepSeek batch token costs by half on the gateway I run — but exact numbers are time- and model-specific, so verify against the pricing page before you budget around them.

CTA: Model roster and live pricing: https://agirouter.org/api/pricing

Disclosure: I'm a founder of AGIRouter. Pricing examples above come from the platform's public pricing API (verified 2026-10-02) and are subject to change; always confirm current rates at agirouter.org/api/pricing. This post was written with AI assistance and fact-checked by the author.

Top comments (0)