DEV Community

Cover image for Which Flash To Run — 五個便宜模型的成本與能力實測
Yang Goufang
Yang Goufang

Posted on

Which Flash To Run — 五個便宜模型的成本與能力實測

一句話摘要: 五個便宜模型沒有總冠軍——日常主力邊際成本零的是 baicodex(qwen3.8-flash),按量最划算且 agentic 四軸全第一的是 opzcode(GLM-5.3-Flash),要 WebSearch 或要快才輪到 deepcode。其餘差別不在牌價,而在到期日、隱形 thinking 與權限的前提。

2026-08-27 實測,主角是五個便宜選項:三家 Flash(Qwen3.8-Flash、GLM-5.3-Flash、DeepSeek V4 Flash Vision-Exp)、Muse Spark 1.2 contributor、Dots3-Note preview:free。先講結論,再攤帳單與數字。

結論:按「你現在要做什麼」選

┌─────────────────────────────────────────────────────────────────────────┐
│  沒有總冠軍,按情境選                                                     │
├──────────────┬──────────────┬──────────────┬────────────────────────────┤
│ 日常主力     │ 按量最佳     │ 跨檔大改     │ 要快 / 要 WebSearch        │
│ baicodex     │ opzcode      │ opzcode      │ deepcode                   │
│ $0           │ $0.05 /M     │ 避開 qwen    │ $0.23 尖峰 · 120 t/s       │
│ qwen3.8-flash│ GLM-5.3 Flash│ GLM 領先 8.2 │ 唯一回 server_tool_use     │
└──────────────┴──────────────┴──────────────┴────────────────────────────┘
Enter fullscreen mode Exit fullscreen mode
情境 用這個 為什麼
日常主力 baicodex(qwen3.8-flash · token plan · Codex) 週配額已付過,邊際成本 $0;agentic 表現與 DeepSeek 同級(四軸各贏兩軸)
按量付費最佳 opzcode(z-ai/glm-5.3-flash:floor · OpenRouter) agentic 四軸全第一、AA 指數並列最高,blended $0.05 只有 DeepSeek 尖峰價的四分之一
跨檔案大改 opzcode Qwen 在 NL2Repo 掉到 48.1,輸 GLM 8.2 分且有使用者回報佐證
要 WebSearch 或要快 deepcode(deepseek-v4-flash-vision-exp) 官方端點唯一回 server_tool_use,輸出 120 t/s 全場最快

命名規則(2026-08-27 起):裸名就是該家族的 flash,旗艦才加標記,標記用廠商自己的字——Qwen 說 max、DeepSeek 說 pro。所以 baicodex 是 flash、baimaxcodex 是旗艦,opgflashcode 改叫 opzcodeflash 不再出現在名字裡。

到期日提醒: opzcode 的 $0.05 到 2026-09-09 為止。Z.ai 在 OpenRouter 的促銷結束,那三家回到兩倍價;而且預設路由不會偏好便宜的那幾家——實測 18 次,裸 id 只有 10 次落在促銷價,加權後 $0.102/M input,比牌價貴 36%。現在已釘 :floor 後綴把它壓回促銷價(20/20 命中)。

成本:牌價不是重點,帳單機制才是

blended 按 7:2:1(cache hit : input : output)估。

成本一覽(每 M tokens)— 純文字,無圖床依賴

baicodex  █ $0.00  ████████████████████ 邊際零成本(週配額已付)
dotscode  █ $0.00  ████████████████████ 免費預覽
musecode  █ $0.041 ████▎
opzcode   █ $0.05  █████▎  (不釘 :floor 實測 $0.068)
deepcode  █ $0.23  ████████████████████████ 尖峰價(週末半價 $0.115)
          └──────────────────────────────
           $0        $0.10        $0.20
Enter fullscreen mode Exit fullscreen mode
Wrapper 牌價 in / out(每 M) Blended 我們實付
baicodex — qwen3.8-flash · token plan · Codex $0.15 / $0.47 $0.088 $0
dotscode — dots-3-note-preview:free $0 / $0 $0 $0
musecode — muse-spark-1.2-contributor $0.10 / $0.20 $0.041 $0.041
opzcode — z-ai/glm-5.3-flash:floor · OpenRouter $0.075 / $0.25 促銷價 2026-09-09 止 $0.05 $0.05(不釘 :floor 實測 $0.068)
deepcode — deepseek-v4-flash-vision-exp 離峰 $0.22 / $0.66 尖峰 $0.44 / $1.32 $0.115 / $0.23 尖峰 $0.23

DeepSeek 的時段差價: 尖峰只在週一至週五 09:00–12:00 與 14:00–18:00(台北時間 = UTC 01–04 / 06–10),週六日全天按離峰計價(2026-08-23 起)。週末用 deepcode 一律半價。

台北時間(週一至週五)— 尖峰 2倍價,週末全離峰
 00         09         12   14         18         24
  ├──────────┼──────────┼────┼──────────┼──────────┤
  │  離峰    │   尖峰   │ 離峰│   尖峰   │   離峰   │
  │ ████████ │ ████████ │ ██ │ ████████ │ ████████ │
  │ 9h       │ 3h       │ 2h │ 4h       │ 6h       │
  └──────────┴──────────┴────┴──────────┴──────────┘
         週六日 ─── 全天離峰(半價)
Enter fullscreen mode Exit fullscreen mode

同一天關掉的門

這些不是表上的選項了,列在這裡是因為名字還會出現在舊筆記和肌肉記憶裡。

  • ocode — 已停用(key 註解掉,row 保留)。Go plan 對每個模型都回 CreditsError,但 qwen3.8-flash 回的是 ModelError: not supported——後者不是帳務症狀,就算續訂也接不到 flash。
  • baicode / baiflashcode — ModelStudio 的 Claude Code wrapper 全數移除。token plan 的 Anthropic 路徑(/apps/anthropic)是沒有內建 server tools 的端點,留著等於讓缺口被誤觸,整家改走 Codex 的 compatible-mode
  • baiflashcodex — 併入 baicodex,刪掉不是留 alias——留 alias 會讓舊肌肉記憶繼續指向貴的那一個。
  • baimaxcodex — 仍在,但在觀察名單上。qwen3.8-max 慢、有幻覺,該淘汰;留著只是佔住旗艦位置等 Qwen 下一版。

能力:廠商自評

表示該廠商沒公布,不是零分。* 標該列最高。

能力對照(廠商自評,非獨立複現)

Terminal Bench 2.1    GLM  84.3*  DeepSeek 82.7  Vision-Exp 83.9  Muse 82.9△  Dots 75.1  Qwen —
DeepSWE v1.1          GLM  63.4*  Qwen 58.7  DeepSeek 54.4
NL2Repo               Vision-Exp 57.7*  GLM 56.3  DeepSeek 54.2  Qwen 48.1
Toolathlon Verified   GLM  78.4*  Qwen 73.5  DeepSeek 70.3
Agents' Last Exam     GLM  26.3*  DeepSeek 25.2  Qwen 24.3
SWE-bench Pro         Qwen 62.5*  DeepSeek 56.0 (*Qwen重測)
AA Index              GLM 57*  Muse 57*  Dots 55.9‡  DeepSeek 52
速度                  DeepSeek 120 t/s*  Vision-Exp 85  Qwen 51  GLM 50
TTFT                  Vision-Exp 0.90s*  DeepSeek 1.05s  GLM 1.47s  Qwen 3.42s
Enter fullscreen mode Exit fullscreen mode
Benchmark GLM-5.3 Flash Qwen3.8 Flash DeepSeek 0731 DeepSeek Vision-Exp Muse 1.2 Dots3 Note
Terminal Bench 2.1 84.3 82.7 83.9 82.9 △ 75.1
DeepSWE v1.1 63.4 58.7 54.4
NL2Repo 56.3 48.1 54.2 57.7
Toolathlon Verified 78.4 73.5 70.3
Agents' Last Exam 26.3 24.3 25.2
SWE-bench Pro 62.5 56.0 *
SWE-bench Verified 78.4
AA Intelligence Index 57 52 57 55.9 ‡
輸出速度 50 t/s 51 t/s 120 t/s 85 t/s
首字延遲 TTFT 1.47s 3.42s 1.05s 0.90s

這些數字不能直接比:

  • △ Muse 的 82.9 跑在 Meta 自家的 Muse Code agent 裡,模型跟 harness 共同訓練;我們的 musecode 是 Claude Code。
  • * SWE-bench Pro 56.0 是 Qwen 自己重測 DeepSeek 的結果,來源不同。
  • ‡ Dots3 的 55.9 來自 ModelCap 指數,不是 AA 指數。
  • Qwen 的空格:Terminal Bench 2.1 至今無公開數字;同代 Max 86.6、27B 73.0,Flash 夾中間是推測所以留空。

牌價看不出來的成本

┌─────────────┬──────────────────────────────────────────────────────────┐
│ musecode    │ Thinking 爆量:say ok 輸入 9 → 輸出 165(154 thinking)  │
│             │ 實付 $0.20/M 在瑣碎工作上行為像 ~$1/M                     │
│             │ Contributor 層拿資料換價格(貴 19 倍才不訓練)           │
├─────────────┼──────────────────────────────────────────────────────────┤
│ dotscode    │ 上游 AtlasCloud content_filter 偶發切斷,無第二家可繞    │
│             │ 免費到 2026-09-30,到期 400                              │
├─────────────┼──────────────────────────────────────────────────────────┤
│ opzcode     │ 12 家只有 3 家給促銷價,不釘 :floor 加權 $0.068          │
│             │ :zzzz 200 照常計費分不出來,只有 @provider 才 400        │
│             │ 2026-09-09 後 :floor 從省錢變風險(262k context 陷阱)  │
├─────────────┼──────────────────────────────────────────────────────────┤
│ 三家權限    │ musecode/dotscode/opzcode 皆 skip_permissions=true       │
│             │ 省錢但整套權限系統關掉(分類器拿不到 claude-opus-5)     │
└─────────────┴──────────────────────────────────────────────────────────┘
Enter fullscreen mode Exit fullscreen mode

musecode — Thinking 爆量,且拿資料換價格。 實測一句 say ok:輸入 9 token,輸出 165,其中 154 是 thinking,且回的是 redacted_thinking 看不到內容。token plan 同題只花 25/34/32。照此比例 $0.20/M 的 output 在瑣碎工作上行為像 ~$1/M。Contributor 層會拿輸入與輸出訓練 Meta 模型,標準層才有不訓練承諾但貴 19 倍。

dotscode — 上游有輸出過濾器。 正常技術對話被 finish_reason: content_filter 切斷;OpenRouter 審核未介入(moderation_latency: null),是唯一 provider AtlasCloud 做的,沒有第二家可繞。以同樣文字重測三次全過,偶發非必現。Free preview 到 2026-09-30,屆時以「未知 id」400。

opzcode — 三個陷阱: 1) 牌價不是實付價,12 家 endpoint 只有 3 家給促銷價(Z.AI、Novita、GMICloud),8 家兩倍,Venice +25%,預設路由不會偏好便宜的——不釘 :floor 實測 $0.068;2) 打錯後綴不會報錯只會多付錢,:zzzz 回 200 照常計費,model 欄位分不出來,只有 @provider 才 400;3) 2026-09-09 後 :floor 會從省錢變風險——它是照價格排序,到期後便宜格會混進 context 只有 262k 的 provider,而此列宣告 1M,宣告過高不會在啟動時報錯,會讓對話在工作中途被拒。

三家都關了權限: musecodedotscodeopzcodeskip_permissions = true——整個權限系統關掉。原因是 auto mode 分類器會去解析第一方的 claude-opus-5,在這些端點上拿不到,每個受管制呼叫都會死。


能力數字全為廠商自評。agentic 四軸(DeepSWE / NL2Repo / Toolathlon / Agents' Last Exam)取自 r/LocalLLM 使用者彙整對照表,已與各廠商公布值交叉核對;Terminal Bench 2.1 的 84.3 與 82.7 為獨立查證吻合。成本與模型回應為 2026-08-27 實測;OpenRouter 實際落點費率由每次回應自己的 cost_details 反推,非牌價。

Top comments (2)

Collapse
 
topstar_ai profile image
Luis Cruz

The analysis of cost versus performance across these models is insightful, especially how you highlight the trade-offs between daily use and peak performance needs. It's fascinating to see how opzcode stands out for use cases requiring adaptable pricing, while deepcode shines for speed. Given the evolving landscape of LLMs, integrating insights from user feedback can further refine model selection for different scenarios. If you're exploring additional support for optimization or implementation in this area, I’d be open to discussing a paid collaboration. What are your thoughts on how user behavior might shift the favorability of these models over time?

Collapse
 
leftoverpzero profile image
Leftover

The 10 out of 18 miss on the promo route is the part I keep. A wrapper name is not a price.

On PZERO leftover daily capacity dies at UTC midnight. I quote the live row when I have a job. If the book is thin I shrink it. Same habit as pinning :floor, just the clock is daily instead of 2026-09-09.