AIHubMix is an AI API gateway — one key and one base URL proxying many vendors' models. MiMo v2.6 Pro is served over an OpenAI-compatible API (callable in the same format as OpenAI's), so the official openai SDK needs only a base_url override.
Key points
- Base URL
https://aihubmix.com/v1, protocol OpenAI Chat Completions - Paid
mimo-v2.6-pro· freexiaomi-mimo-v2.6-pro-free[1] [2] - $0.48/M input · $0.96/M output · $0.0038/M cache read [1]
- Free tier 5 req/min · 100 req/day · 1M tokens/day [2]
- Context 1.05M tokens · max output 131K tokens [1]
- Dependencies:
openai>=1.40,python-dotenv>=1.0
Specifications
| Base URL | https://aihubmix.com/v1 |
| Protocol | OpenAI Chat Completions |
| Paid model ID | mimo-v2.6-pro |
| Free model ID | xiaomi-mimo-v2.6-pro-free |
| Other variants |
mimo-v2.6-pro-ultraspeed, mimo-v2.6-flash, mimo-v2.5-pro, coding-xiaomi-mimo-v2.6-pro [3] |
| Pricing (paid) | $0.48/M input · $0.96/M output · $0.0038/M cache read [1] |
| Free tier limits | 5 req/min · 100 req/day · 1M tokens/day [2] |
| Context window | 1.05M tokens [1] |
| Max output | 131K tokens [1] |
| Input modalities | Text, vision, audio, video — text output only [1] |
| Listed features | Streaming, tool calling, structured outputs, prompt caching [1] |
| Tested with | Python 3.14.5, openai 3.19.2 |
Note that the paid and free tiers use different naming conventions: free variants carry the xiaomi- prefix, paid ones do not.
Steps
1. Environment.
python3 -m venv .venv
.venv/bin/pip install "openai>=1.40" "python-dotenv>=1.0"
2. Credentials. Keep the key out of source and out of git.
printf 'AIHUBMIX_API_KEY=your_key_here\n' > .env
chmod 600 .env
printf '.env\n.venv/\n__pycache__/\n' > .gitignore
3. Client. chat.py:
import argparse
import os
import sys
from dotenv import load_dotenv
from openai import (
APIConnectionError,
APIStatusError,
AuthenticationError,
OpenAI,
RateLimitError,
)
BASE_URL = "https://aihubmix.com/v1"
MODEL = "mimo-v2.6-pro"
def build_client() -> OpenAI:
load_dotenv()
api_key = os.environ.get("AIHUBMIX_API_KEY", "")
if not api_key or api_key.startswith("<"):
sys.exit("AIHUBMIX_API_KEY not found — set it in .env")
return OpenAI(api_key=api_key, base_url=BASE_URL)
def _msg(exc: APIStatusError) -> str:
try:
bodies = (exc.body, exc.response.json())
except Exception:
bodies = (exc.body,)
for body in bodies:
if isinstance(body, dict) and isinstance(body.get("error"), dict):
if body["error"].get("message"):
return str(body["error"]["message"])
return str(exc)
def run(client: OpenAI, args, messages: list) -> None:
if args.stream:
stream = client.chat.completions.create(
model=args.model, messages=messages,
max_tokens=args.max_tokens, stream=True,
)
for chunk in stream:
if not chunk.choices:
continue
piece = chunk.choices[0].delta.content
if piece:
print(piece, end="", flush=True)
print()
else:
response = client.chat.completions.create(
model=args.model, messages=messages, max_tokens=args.max_tokens,
)
print(response.choices[0].message.content)
if response.usage:
print(f"\n[tokens] {response.usage.total_tokens}", file=sys.stderr)
def main() -> None:
parser = argparse.ArgumentParser(description="Call Xiaomi MiMo via AIHubMix")
parser.add_argument("prompt", nargs="?", default="Introduce yourself in one sentence.")
parser.add_argument("--model", default=MODEL)
parser.add_argument("--max-tokens", type=int, default=1024)
parser.add_argument("--no-stream", dest="stream", action="store_false")
args = parser.parse_args()
client = build_client()
messages = [{"role": "user", "content": args.prompt}]
try:
run(client, args, messages)
except AuthenticationError:
sys.exit("Auth failed: AIHUBMIX_API_KEY invalid or revoked.")
except RateLimitError as exc:
sys.exit(f"Rate limited (429): {_msg(exc)}")
except APIConnectionError as exc:
sys.exit(f"Connection failed: {exc}")
except APIStatusError as exc:
sys.exit(f"API returned {exc.status_code}: {_msg(exc)}")
if __name__ == "__main__":
main()
4. Run.
.venv/bin/python chat.py "Introduce yourself in one sentence."
I'm MiMo, a large language model developed by Xiaomi's LLM Core Team, here to assist with answering questions, writing, and a wide range of helpful tasks.
Roughly 6s to first tokens in my runs, against the 6.4s listed on the model page [1].
Implementation notes
Streaming loop guards. Both if statements are load-bearing. Chunks with an empty choices array do arrive, and chunk.choices[0] raises IndexError on them. delta.content is None on role-announcement and finish frames, which print renders as the literal string None. flush=True is what makes output stream rather than appear in buffered blocks.
Error message extraction. Reading exc.body alone is not reliable — it can fall through to str(exc), which prints the raw Error code: 429 - {'error': ...} dict repr. Parsing exc.response.json() as a fallback recovers the actual message. That difference matters for the 429 case below, where the message text is the only thing distinguishing two conditions.
Placeholder guard. api_key.startswith("<") catches an unedited <YOUR_AIHUBMIX_API_KEY>, which otherwise reaches the API and returns a 401 you have to interpret.
Usage
.venv/bin/python chat.py "your prompt" # stream, default model
.venv/bin/python chat.py --no-stream "your prompt" # buffered + token count
.venv/bin/python chat.py --model mimo-v2.6-flash "..." # override model
.venv/bin/python chat.py --max-tokens 4096 "..." # raise output cap
Output truncated mid-sentence is your --max-tokens cap, not a model limit.
Troubleshooting
| Symptom | Cause | Action |
|---|---|---|
429 "rate limited by provider" |
Per-minute throttle | Back off and retry |
429 "reached the limit of the free model quota" |
Account free quota exhausted, shared across free models | Retry cannot succeed; switch to mimo-v2.6-pro
|
400 "cannot be served at the moment" |
No upstream channel; see no_available_channel
|
Provider-side; wait or contact support |
401 |
Bad or revoked key | Check .env
|
Branch retry logic on the message body, not the status code — throttling and quota exhaustion share 429 but need opposite handling.
Confirm a model ID against the API rather than documentation:
curl -s https://aihubmix.com/v1/models \
-H "Authorization: Bearer $AIHUBMIX_API_KEY" \
| python -c "
import json,sys
ids=[m['id'] for m in json.load(sys.stdin)['data']]
print('exact match:', 'mimo-v2.6-pro' in ids)
"
The SDK's exception rendering drops the machine-readable error code. Reproduce with curl to see it:
curl -s https://aihubmix.com/v1/chat/completions \
-H "Authorization: Bearer $AIHUBMIX_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"mimo-v2.6-pro","messages":[{"role":"user","content":"hi"}],"max_tokens":30}'
A body containing "code": "no_available_channel" means the ID is valid but unrouteable at that moment. I observed this on all paid mimo variants; the same configuration succeeded later with no local changes.
Attributing a failure. Call a different vendor's model with the same key and script:
.venv/bin/python chat.py --no-stream --max-tokens 30 --model gpt-4o-mini "reply: ok"
Success clears your key, base URL, request format and network path in one call, leaving the specific model as the only remaining variable. Failure points at your setup or account instead. Roughly 12 tokens.
Billing can be ruled out directly:
curl -s https://aihubmix.com/v1/dashboard/billing/subscription \
-H "Authorization: Bearer $AIHUBMIX_API_KEY"
Check has_payment_method and the limit fields.
Sources
- [1] AIHubMix model page,
mimo-v2.6-pro: https://aihubmix.com/model/mimo-v2.6-pro (accessed 28 September 2026) - [2] AIHubMix model page,
xiaomi-mimo-v2.6-pro-free: https://aihubmix.com/model/xiaomi-mimo-v2.6-pro-free (accessed 28 September 2026) - [3] Variant IDs enumerated from the
GET /v1/modelsendpoint (accessed 28 September 2026)
Top comments (0)