<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: sunMM</title>
    <description>The latest articles on DEV Community by sunMM (@iceseaboy).</description>
    <link>https://dev.to/iceseaboy</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4123273%2F29f1ab61-cade-4a50-9a16-f98b6b6977ff.jpg</url>
      <title>DEV Community: sunMM</title>
      <link>https://dev.to/iceseaboy</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/iceseaboy"/>
    <language>en</language>
    <item>
      <title>Run Claude Code on MiniMax-M3 with a 1M context for about $0.30 per million input tokens</title>
      <dc:creator>sunMM</dc:creator>
      <pubDate>Sun, 13 Sep 2026 14:40:23 +0000</pubDate>
      <link>https://dev.to/iceseaboy/run-claude-code-on-minimax-m3-with-a-1m-context-for-about-030-per-million-input-tokens-2n30</link>
      <guid>https://dev.to/iceseaboy/run-claude-code-on-minimax-m3-with-a-1m-context-for-about-030-per-million-input-tokens-2n30</guid>
      <description>&lt;p&gt;If you use Claude Code every day, you already know the shape of the bill. Agentic coding is not one prompt; it is hundreds of &lt;code&gt;/v1/messages&lt;/code&gt; calls per session, each one re-sending the system prompt, the tool definitions, the file contents you just read, and the whole conversation so far. Input tokens dominate. A long refactor across a large repo can burn through more context in an afternoon than a chat user sends in a month.&lt;/p&gt;

&lt;p&gt;Claude Code itself is not the expensive part. The model behind it is. And Claude Code will happily talk to any server that speaks the Anthropic Messages API, which means you can keep the workflow you like and swap the model when the task does not need a frontier model.&lt;/p&gt;

&lt;p&gt;This post shows how to point Claude Code at MiniMax-M3 or MiniMax-M2.7 through an Anthropic-compatible relay, what it costs, and where it falls short. No benchmarks, no testimonials; just the setup and the caveats.&lt;/p&gt;

&lt;h2&gt;
  
  
  What MiniMax-M3 and M2.7 are
&lt;/h2&gt;

&lt;p&gt;MiniMax is a Chinese lab whose M-series models are built for agentic and coding work. Two of them matter here:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;MiniMax-M2.7&lt;/strong&gt;: the everyday coding model. Edits, refactors, tool calls, test loops.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MiniMax-M3&lt;/strong&gt;: same per-token price as M2.7 up to 512K tokens, with a &lt;strong&gt;1,048,576-token context window&lt;/strong&gt;. That is the one you reach for when you want a whole repository, or a very long agent session, in a single context.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There is also &lt;strong&gt;MiniMax-M2.7-highspeed&lt;/strong&gt;, a faster variant priced higher per token, which is a good fit for Claude Code's "small fast model" slot (the one it uses for quick helper calls).&lt;/p&gt;

&lt;p&gt;The relay in this guide is &lt;a href="https://yiduochan.com" rel="noopener noreferrer"&gt;YiduoChan&lt;/a&gt;, a gateway built on the open-source new-api project. It exposes these models on two surfaces:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;OpenAI-compatible: &lt;code&gt;https://yiduochan.com/v1&lt;/code&gt; (chat completions, audio, video)&lt;/li&gt;
&lt;li&gt;Anthropic-compatible: &lt;code&gt;https://yiduochan.com/v1/messages&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Claude Code uses the second one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Setup: environment variables
&lt;/h2&gt;

&lt;p&gt;Create an account at &lt;code&gt;https://yiduochan.com/register&lt;/code&gt;, generate an API key, then export four variables before you launch Claude Code:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;ANTHROPIC_BASE_URL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;https://yiduochan.com
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;ANTHROPIC_AUTH_TOKEN&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;YIDUOCHAN_API_KEY
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;ANTHROPIC_MODEL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;MiniMax-M2.7
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;ANTHROPIC_SMALL_FAST_MODEL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;MiniMax-M2.7-highspeed

claude
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Swap &lt;code&gt;ANTHROPIC_MODEL=MiniMax-M2.7&lt;/code&gt; for &lt;code&gt;ANTHROPIC_MODEL=MiniMax-M3&lt;/code&gt; when you want the 1M context. Note that &lt;code&gt;ANTHROPIC_BASE_URL&lt;/code&gt; is the bare host; Claude Code appends &lt;code&gt;/v1/messages&lt;/code&gt; itself.&lt;/p&gt;

&lt;p&gt;If you want to flip between Anthropic and MiniMax per shell, a tiny function works well:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;cc-minimax&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
  &lt;span class="nv"&gt;ANTHROPIC_BASE_URL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;https://yiduochan.com &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nv"&gt;ANTHROPIC_AUTH_TOKEN&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$YIDUOCHAN_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nv"&gt;ANTHROPIC_MODEL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;MiniMax-M3 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nv"&gt;ANTHROPIC_SMALL_FAST_MODEL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;MiniMax-M2.7-highspeed &lt;span class="se"&gt;\&lt;/span&gt;
  claude &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$@&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Setup: settings.json
&lt;/h2&gt;

&lt;p&gt;For something persistent, put the same values in &lt;code&gt;~/.claude/settings.json&lt;/code&gt; (global) or &lt;code&gt;.claude/settings.json&lt;/code&gt; inside a project (so only that repo uses MiniMax):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"env"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"ANTHROPIC_BASE_URL"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://yiduochan.com"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"ANTHROPIC_AUTH_TOKEN"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"YIDUOCHAN_API_KEY"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"ANTHROPIC_MODEL"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"MiniMax-M3"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"ANTHROPIC_SMALL_FAST_MODEL"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"MiniMax-M2.7-highspeed"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The per-project variant is the one I would recommend: keep your real Claude setup as the default and opt a specific repo into the cheaper model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verify the endpoint with curl
&lt;/h2&gt;

&lt;p&gt;Before blaming Claude Code for anything, confirm the relay answers a plain Messages request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl https://yiduochan.com/v1/messages &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"x-api-key: &lt;/span&gt;&lt;span class="nv"&gt;$YIDUOCHAN_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"anthropic-version: 2023-06-01"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"model": "MiniMax-M2.7", "max_tokens": 64,
       "messages": [{"role": "user", "content": "Say hello"}]}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You should get back a standard Anthropic Messages response (a &lt;code&gt;content&lt;/code&gt; array plus a &lt;code&gt;usage&lt;/code&gt; block). Without a key the same route returns HTTP 401, so if you see that, the header is the problem, not the model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pricing
&lt;/h2&gt;

&lt;p&gt;All prices are in USD per 1M tokens and come from the relay's public pricing page at the time of writing. They are MiniMax's list prices converted at a fixed exchange rate rather than a marked-up subscription.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Input&lt;/th&gt;
&lt;th&gt;Output&lt;/th&gt;
&lt;th&gt;Cache read&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;MiniMax-M3 (up to 512K)&lt;/td&gt;
&lt;td&gt;$0.30&lt;/td&gt;
&lt;td&gt;$1.20&lt;/td&gt;
&lt;td&gt;$0.06&lt;/td&gt;
&lt;td&gt;1M context&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MiniMax-M3 (above 512K)&lt;/td&gt;
&lt;td&gt;$0.60&lt;/td&gt;
&lt;td&gt;$2.40&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;tier kicks in past 512K&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MiniMax-M2.7&lt;/td&gt;
&lt;td&gt;$0.30&lt;/td&gt;
&lt;td&gt;$1.20&lt;/td&gt;
&lt;td&gt;$0.06&lt;/td&gt;
&lt;td&gt;cache write $0.375&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MiniMax-M2.7-highspeed&lt;/td&gt;
&lt;td&gt;$0.60&lt;/td&gt;
&lt;td&gt;$2.40&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;small fast model&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For reference, MiniMax's own international pricing for M2.7 is about $0.30 in / $1.20 out, so the relay is roughly at parity with going direct; the point is the Anthropic-compatible endpoint plus the M3 tier, not a discount on MiniMax itself.&lt;/p&gt;

&lt;p&gt;Two things to notice for Claude Code specifically:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Cache reads at $0.06&lt;/strong&gt; matter a lot, because Claude Code re-sends the same system prompt and tool schema on every turn. Prompt caching is applied automatically on the relay side.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The 512K tier boundary on M3&lt;/strong&gt; is real. If you routinely push past half a million tokens of context, your marginal cost doubles. Use &lt;code&gt;/compact&lt;/code&gt; before you get there unless you genuinely need the whole repo resident.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Trying it without prepaying anything
&lt;/h2&gt;

&lt;p&gt;The obvious objection to any small relay is that you have to wire it money first.&lt;br&gt;
I didn't want to ask for that, so there's a per-call endpoint that takes payment&lt;br&gt;
on the request itself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST https://yiduochan.com/api/x402/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"model":"MiniMax-M2.7","messages":[{"role":"user","content":"ping"}]}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That first call comes back as &lt;code&gt;402 Payment Required&lt;/code&gt; with the terms in a header:&lt;br&gt;
one cent in USDC, on Base or Arbitrum. Any x402 client signs it and retries, and&lt;br&gt;
the second response body is the completion. No account, no email, no balance to&lt;br&gt;
lose. The API key comes back in a header afterwards, so once you're satisfied it&lt;br&gt;
works you can drop back to the normal &lt;code&gt;ANTHROPIC_AUTH_TOKEN&lt;/code&gt; flow above.&lt;/p&gt;

&lt;p&gt;If you'd rather just hold a balance, the credit tiers start at ten cents. Either&lt;br&gt;
way the point is that you get to check the thing works before it holds any real&lt;br&gt;
money of yours.&lt;/p&gt;

&lt;h2&gt;
  
  
  Honest caveats
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;It is not Claude.&lt;/strong&gt; MiniMax-M3 and M2.7 are different models with different training and different habits around tool use. Expect differences in how it plans multi-file changes, how it formats edits and how often it retries a tool call. Run it on your own codebase for a day and decide for yourself; I am deliberately not quoting benchmark numbers here, and you should be suspicious of anyone who does without showing their harness.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reasoning shows up as &lt;code&gt;&amp;lt;think&amp;gt;&lt;/code&gt; tags on the OpenAI endpoint.&lt;/strong&gt; If you use the OpenAI-compatible &lt;code&gt;/v1/chat/completions&lt;/code&gt; surface from Cursor, Continue or your own scripts, the model's reasoning can arrive inline in &lt;code&gt;content&lt;/code&gt; wrapped in &lt;code&gt;&amp;lt;think&amp;gt;...&amp;lt;/think&amp;gt;&lt;/code&gt;. Strip it before you show it to users. Claude Code uses the Messages endpoint, so this does not affect the setup above, but it bites people who mix both surfaces.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;/v1/messages/count_tokens&lt;/code&gt; is not supported.&lt;/strong&gt; The relay currently returns HTTP 404 for &lt;code&gt;POST /v1/messages/count_tokens&lt;/code&gt;; the route is disabled in the gateway code. Anything that relies on server-side token counting will fail. Claude Code's chat and tool loop only needs &lt;code&gt;/v1/messages&lt;/code&gt;, but if you have scripts that call &lt;code&gt;count_tokens&lt;/code&gt;, estimate locally instead.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It is a relay.&lt;/strong&gt; Your prompts, file contents and tool results pass through a third-party server on their way to MiniMax. Treat it exactly as you would any other hosted API: no production secrets in context, no code you are not allowed to send to a third party.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prices can change.&lt;/strong&gt; The numbers above are what the pricing page showed when I wrote this. Check &lt;code&gt;https://yiduochan.com/pricing&lt;/code&gt; before you plan a budget around them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrapping up
&lt;/h2&gt;

&lt;p&gt;Three environment variables (four if you set the small fast model) turn Claude Code into a MiniMax client. Use M2.7 for everyday work, M3 when you want a million tokens of context, and keep your Claude configuration around for the problems that still need it.&lt;/p&gt;

&lt;p&gt;The full guide, including notes for Cline, Roo Code and the Anthropic SDKs, lives at &lt;strong&gt;&lt;a href="https://yiduochan.com/minimax/claude-code/" rel="noopener noreferrer"&gt;https://yiduochan.com/minimax/claude-code/&lt;/a&gt;&lt;/strong&gt;.&lt;/p&gt;




</description>
      <category>ai</category>
      <category>llm</category>
      <category>tutorial</category>
      <category>devtools</category>
    </item>
  </channel>
</rss>
