DEV Community

Greta Volkov
Greta Volkov

Posted on

SiliconFlow API Review 2026: Setup, Models and Real Pricing

Verdict up front: SiliconFlow is a solid, OpenAI-compatible way to call open-weight LLMs (DeepSeek, Qwen, GLM, Kimi, MiniMax) from one key, with an Anthropic-compatible endpoint as a bonus. It is not automatically the cheapest host for every model. Check the specific model you need before you commit. Prices below were read off each provider's public pricing page on October 6, 2026.

I went looking for a current, practical write-up on SiliconFlow and mostly found launch posts. SiliconFlow's own dev.to account announces models as they arrive, which is useful, but its two newest posts there cover GLM-4.6V and GLM-4.7, and both have since been deprecated on the platform. So this is the guide I wanted: how to connect, how to find the model IDs that actually exist today, what it costs compared with the alternatives, and where it falls short.

One limit to know early: SiliconFlow is mostly a text and LLM platform. Its pricing page lists a handful of FLUX image models, two Wan 2.2 video models and two TTS models. For image, video and speech work I use synexa instead, which runs 100+ media models behind a single predictions endpoint. The rest of this post is about the LLM side, where SiliconFlow is strongest.

1. Connecting to the SiliconFlow API

Setup is the standard OpenAI-compatible routine. Create an account, open API Keys in the console, generate a key, and point any OpenAI SDK at https://api.siliconflow.com/v1. New accounts get $1 in free credits, and billing after that is pay-as-you-go with no minimum commitment.

The first thing I do with any new provider is list the models through the API instead of trusting the marketing page:

export SILICONFLOW_API_KEY="sk-..."

# Chat models only; other sub_type values: embedding, reranker,
# text-to-image, image-to-image, speech-to-text, text-to-video
curl -s "https://api.siliconflow.com/v1/models?sub_type=chat" \
  -H "Authorization: Bearer $SILICONFLOW_API_KEY" | jq -r '.data[].id'
Enter fullscreen mode Exit fullscreen mode

Then a minimal Python call with the official openai package:

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["SILICONFLOW_API_KEY"],
    base_url="https://api.siliconflow.com/v1",
)

resp = client.chat.completions.create(
    model="deepseek-ai/DeepSeek-V4.1-Flash",
    messages=[{"role": "user", "content": "Summarise this changelog in 3 bullets: ..."}],
    max_tokens=512,
)
print(resp.choices[0].message.content)
print(resp.usage)  # log this; it's what you're billed on
Enter fullscreen mode Exit fullscreen mode

Model IDs use the org/Model form (deepseek-ai/..., zai-org/..., moonshotai/..., Qwen/...), so copy them exactly from the /models output.

2. Why /models beats the API reference

This is the most useful thing I found. The chat completions reference documents the model parameter with an enum of allowed values, and that list is out of date. It still includes zai-org/GLM-4.7, zai-org/GLM-4.6 and nex-agi/DeepSeek-V3.1-Nex-N1, which the release notes say were deprecated in May and June 2026. It doesn't include deepseek-ai/DeepSeek-V4.1-Flash, which has a live model page and price.

So treat the reference as a guide to parameters, not to models. Two parameters worth knowing about:

  • enable_thinking switches thinking mode on or off, but only for the models listed in the reference. Since it isn't a standard OpenAI field, pass it via extra_body in the Python SDK.
  • thinking_budget caps chain-of-thought tokens for reasoning models (minimum 128, maximum 32,768, default 4,096).

3. Deprecations come with traffic migration

SiliconFlow retires models on a regular schedule and posts notices in its release notes. Sometimes it also reroutes traffic. On June 11, 2026, calls to zai-org/GLM-5 were automatically routed to zai-org/GLM-5.1, and moonshotai/Kimi-K2.5 to Kimi-K2.6, at the same price.

That's convenient, but it means the model answering your request can change while your code stays the same. If you run evals or care about reproducible output, pin the successor name yourself and watch the release notes.

4. SiliconFlow pricing compared with alternatives

Prices are per 1M tokens, shown as input / cached input / output. A dash means I couldn't find the model on that provider's pricing page.

Model SiliconFlow DeepSeek direct Together AI DeepInfra
DeepSeek V4.1 Flash $0.15 / $0.003 / $0.60 $0.15 / $0.003 / $0.60 off-peak; $0.30 / $0.006 / $1.20 peak $0.30 / $0.006 / $1.20 —
DeepSeek V4 Pro 0813 $1.32 / $0.044 / $3.96 $0.66 / $0.022 / $1.98 off-peak; $1.32 / $0.044 / $3.96 peak $1.32 / $0.13 / $3.96 —
DeepSeek V4 Flash 0731 $0.22 / $0.014 / $0.66 retired, served by V4.1 Flash $0.14 / $0.03 / $0.28 $0.06 / $0.015 / $0.18
GLM-5.3 $1.40 / $0.26 / $4.40 — $1.40 / $0.26 / $4.40 —
Kimi K3 $2.70 / $0.27 / $13.50 — $2.70 / $0.27 / $13.50 $2.85 / $0.285 / $14.25
Kimi K2.6 $0.77 / $0.14 / $3.40 — — $0.75 / $0.15 / $3.50
MiniMax M3 $0.30 / $0.06 / $1.20 — $0.30 / $0.06 / $1.20 —

What the table actually says:

  • DeepSeek V4.1 Flash is SiliconFlow's best deal. You pay DeepSeek's off-peak rate at all hours. DeepSeek's own peak windows (01:00–04:00 and 06:00–10:00 UTC, weekdays) cost double, and Together charges the peak rate all the time.
  • DeepSeek V4 Pro 0813 is the opposite. SiliconFlow charges DeepSeek's peak rate at all hours. If your jobs can run off-peak, going direct costs half.
  • Older revisions can cost much more. For V4 Flash 0731, SiliconFlow's input and output prices are about 3.7x DeepInfra's.
  • Mainstream models are at parity. GLM-5.3, Kimi K3 and MiniMax M3 match Together's prices to the cent, so for those the decision comes down to rate limits and reliability, not price.

If you'd rather keep several providers behind one key, OpenRouter says it passes provider prices through without markup and charges a 5.5% fee ($0.80 minimum) when you buy credits.

5. Rate limits start low for long-context work

Limits are per account, not per key, and they apply per model. Hitting the limit on one model doesn't throttle the others. Your tier depends on monthly spend (the higher of last month and month-to-date) and upgrades automatically:

Tier RPM TPM
L0 (new accounts) 1,000 40,000
L2 2,000 80,000
L4 8,000 500,000
L5 10,000 2,000,000

The catch: many current models advertise context windows of around 1M tokens, but an L0 account gets 40,000 tokens per minute. If you plan to send whole repositories or long documents, budget for spend that moves you up a tier, or test early whether you'll hit 429s.

6. Bonus: using SiliconFlow with Claude Code

SiliconFlow also exposes an Anthropic-style Messages endpoint, and its docs include a Claude Code integration. The manual version is three environment variables:

export ANTHROPIC_BASE_URL="https://api.siliconflow.com/"
export ANTHROPIC_MODEL="your-preferred-model"   # a model ID from /models
export ANTHROPIC_API_KEY="sk-your-siliconflow-key"
Enter fullscreen mode Exit fullscreen mode

The docs also offer a one-line setup script piped from a URL. Read it before you run it; the docs say so too.

Who should use SiliconFlow

Good fit: you want many open-weight LLMs behind one OpenAI-compatible key, your traffic peaks during DeepSeek's peak hours, or you want Claude Code-style tooling on top of open models.

Look elsewhere (or compare per model): you're cost-sensitive on older model revisions, you can batch DeepSeek Pro jobs off-peak, or your workload is mainly image and video generation.

FAQ

Is SiliconFlow OpenAI-compatible?
Yes. Use base URL https://api.siliconflow.com/v1 with the OpenAI SDK. It supports most OpenAI parameters, plus its own fields like enable_thinking and thinking_budget.

Does SiliconFlow have a free tier?
New accounts get $1 in free credits. After that it's pay-as-you-go, with no minimum commitment.

What are SiliconFlow's rate limits?
They're tiered by monthly spend, starting at 1,000 RPM and 40,000 TPM for new accounts and going up to 10,000 RPM and 2,000,000 TPM at L5. Limits apply per account and per model.

Is SiliconFlow cheaper than calling DeepSeek directly?
For DeepSeek V4.1 Flash during DeepSeek's peak hours, yes. Off-peak it costs the same. For V4 Pro 0813, going direct off-peak costs half as much.

Wrapping up

SiliconFlow earns its place as a single endpoint for current open-weight LLMs, and the official quickstart gets you to a first call in minutes. Just don't assume it's cheapest on every model: list models through the API, pin model names, and compare each model's price before you commit. For the media side of my projects (images, video, speech) I keep using synexa, where the same /v1/predictions request shape works across its model gallery.

Top comments (0)