DEV Community

Julia
Julia

Posted on

How to Call 4 Free LLM Models from One OpenAI-Compatible Endpoint (Tested October 2026)

How to Call 4 Free LLM Models from One OpenAI-Compatible Endpoint (Tested October 2026)

Every "free LLM API" list looks great until you integrate it. Then the 502s start: a model that worked last week 404s this week, a provider silently swaps you to a weaker version, and switching upstreams means editing your app code again.

I got tired of re-integrating. The fix I landed on: one OpenAI-compatible endpoint in front of several daily-tested free models, so swapping upstreams is a model string change — not a refactor.

Here is the exact setup, with 4 free models I actually run in production this month. Every example is copy-paste runnable.

Step 0 — One key, one endpoint

Grab a key from apishare.cc. The gateway speaks the standard /v1/chat/completions protocol, so if your code already talks to OpenAI, zero code changes are needed:

export APISHARE_KEY="your-key-here"
Enter fullscreen mode Exit fullscreen mode

All 4 models below use the same base URL and the same key. Only the model field changes.

Model 1: qwen3.8-max — the workhorse

Full ID: ApiShare/alibaba/qwen3.8-max:free. This is what I reach for when the task is open-ended: summarization, refactoring suggestions, long-document Q&A.

curl https://apishare.cc/v1/chat/completions \
  -H "Authorization: Bearer $APISHARE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "ApiShare/alibaba/qwen3.8-max:free",
    "messages": [{"role": "user", "content": "Summarize the tradeoffs of event-driven vs request-response architectures in 5 bullets."}]
  }'
Enter fullscreen mode Exit fullscreen mode

Use it for: general chat, long-context reasoning, anything where answer quality matters more than latency.

Model 2: Hunyuan-MT-7B — translation specialist

Full ID: JamesFoster/tencent/Hunyuan-MT-7B. A 7B machine-translation-focused model. Small, fast, and noticeably better at preserving tone in translation than generic chat models of the same size.

curl https://apishare.cc/v1/chat/completions \
  -H "Authorization: Bearer $APISHARE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "JamesFoster/tencent/Hunyuan-MT-7B",
    "messages": [{"role": "user", "content": "Translate to German: The deployment pipeline is green again after we pinned the base image."}]
  }'
Enter fullscreen mode Exit fullscreen mode

Use it for: localization pipelines, bilingual side projects, batch document translation.

Model 3: GLM-Z1-9B — fast reasoning on a budget

Full ID: MeiLin/THUDM/GLM-Z1-9B-0414. When you need structured output or light chain-of-thought but can't burn 30 seconds waiting, this 9B reasoning model is the sweet spot: fast enough for interactive loops, cheap enough (free) to call in a tight retry loop.

curl https://apishare.cc/v1/chat/completions \
  -H "Authorization: Bearer $APISHARE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "MeiLin/THUDM/GLM-Z1-9B-0414",
    "messages": [{"role": "user", "content": "Extract JSON: invoice #4471 dated 2026-10-01 for $240 from Acme Corp."}],
    "response_format": {"type": "json_object"}
  }'
Enter fullscreen mode Exit fullscreen mode

Use it for: extraction, classification, tool-call argument generation, MVP prototypes.

Model 4: agnes-3.0-flash — the most-called free model

Full ID: Agnes/agnes-3.0-flash:free. This one is the platform's most-called free model (4.3K+ calls and climbing), and it supports streaming plus image input — handy for quick vision tasks like reading a screenshot.

curl https://apishare.cc/v1/chat/completions \
  -H "Authorization: Bearer $APISHARE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Agnes/agnes-3.0-flash:free",
    "messages": [
      {"role": "user", "content": "Describe what this UI screenshot shows in 2 sentences."}
    ]
  }'
Enter fullscreen mode Exit fullscreen mode

Use it for: high-frequency chat, vision quick-checks, anywhere you'd otherwise pay per token.

The billing answer nobody gives you: per-call cost in the response

One thing I wish more providers did: every response comes back with a usage block that states the actual cost:

"apishare_usage": {
  "inputTokens": 79,
  "outputTokens": 2,
  "totalCost": 0
}
Enter fullscreen mode Exit fullscreen mode

totalCost: 0, every call. That's the "Honest Billing" design: the platform meters what you use and shows you the number instead of hiding it behind a dashboard. When one of these models eventually moves to a paid tier, you'll see the cost trail before it becomes a surprise invoice — no silent charges, and never a claim of "unlimited" that quietly has conditions attached.

How I pick which model is alive today

Free tiers die fast, so I don't trust announcements — I trust daily measurements. The Free LLM API Rankings on apishare.cc re-test availability, latency, and throughput every day, and the results are published without sponsored placements ("no sponsored rankings" is a hard rule there — placements never buy a better rank).

Current snapshot (October 2026): 8 free models live, 12K+ total calls served, 88 registered users. The gateway also exposes a first-party official-sources catalog — apishare's own free promo models (like agnes-3.0-flash:free) are listed with an explicit provider field so you always know who's serving your traffic.

The pattern that survives

The whole point is that your app code should never know or care which upstream is behind a model ID today:

  1. One OpenAI-compatible endpoint + one key.
  2. A model string per task (workhorse / translation / reasoning / vision).
  3. Daily-tested availability data instead of vendor marketing.
  4. Per-call cost visibility from day one.

Swapping an upstream when a free tier dies takes me about 30 seconds — one string in config, zero redeploy of application logic.


Disclosure: I run apishare.cc, the gateway and daily-tested ranking described above. This is a partially promotional post, written to the best of my knowledge with real measured numbers. If you're building against free LLM APIs, I'd genuinely like to hear which upstreams have been most stable for you lately.

Top comments (0)