DEV Community

Cover image for A Practical Guide to Building with the Kimi K3 Chat Completions API
Germey
Germey

Posted on Originally published at platform.acedata.cloud

A Practical Guide to Building with the Kimi K3 Chat Completions API

Reasoning models are useful only when your application can call them predictably, stream partial output, and preserve enough conversation state for follow-up turns.

This guide walks through a small, practical Kimi K3 chat-completion workflow using Ace Data Cloud's Kimi endpoint. We will cover the request shape, the response fields you should actually read, how to enable streaming, and how to pass multi-turn messages without inventing a custom protocol.

What you can do

The Kimi Chat Completion API lets you call the kimi-k3 model through an HTTP API. The documented use cases include ordinary chat completion, streaming responses, multi-turn dialogue, and K3 reasoning intensity control through reasoning_effort.

The important request fields are:

Field Purpose
model Selects the Kimi model. The guide recommends kimi-k3.
messages An array of dialogue messages. Each item has role and content.
role Supports user, assistant, system, and tool.
reasoning_effort Top-level field for K3 reasoning. The supported value is max.
stream Set to true when you want line-by-line streaming output.

The endpoint used throughout the guide is:

POST https://api.acedata.cloud/kimi/chat/completions
Enter fullscreen mode Exit fullscreen mode

Authentication is sent with a bearer token:

Authorization: Bearer $ACEDATACLOUD_API_KEY
Enter fullscreen mode Exit fullscreen mode

How it works

At the simplest level, you send a JSON body containing the model and a messages array. Kimi returns a Chat Completions-style response with an id, model, choices, and usage.

Here is a minimal request that asks Kimi K3 to review code and provide a fix:

curl https://api.acedata.cloud/kimi/chat/completions \
  -H "Authorization: Bearer $ACEDATACLOUD_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "kimi-k3",
    "messages": [
      {"role": "user", "content": "Review this code and provide a fix"}
    ],
    "reasoning_effort": "max"
  }'
Enter fullscreen mode Exit fullscreen mode

A normal response contains a choices array. The assistant reply is in choices[0].message.content. The usage object reports token counts, including prompt_tokens, completion_tokens, and total_tokens.

A shortened response looks like this:

{
  "id": "msg_2D4Btbg1WgvkNE3tCYkR4xGA",
  "object": "chat.completion",
  "model": "kimi-k3",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Hello! How can I help you today?"
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 86,
    "completion_tokens": 206,
    "total_tokens": 292
  }
}
Enter fullscreen mode Exit fullscreen mode

For a first integration, store the returned id for observability, read choices[0].message, and log usage so you can understand how prompts grow over time.

Use reasoning_effort carefully

For kimi-k3, reasoning is always enabled. The documented top-level request field is reasoning_effort, and the supported value is currently max. If you omit the field, the behavior is also max.

That means you should not build application logic that depends on unsupported strings such as standard or high. They may be partially accepted by upstream compatibility layers, but the guide explicitly says not to rely on them changing reasoning behavior.

In Python with an OpenAI-style client, the field can be passed directly:

response = client.chat.completions.create(
    model="kimi-k3",
    messages=[{"role": "user", "content": "Design a reliable task queue"}],
    reasoning_effort="max",
)
Enter fullscreen mode Exit fullscreen mode

For multi-turn dialogues and tool calls, return the complete assistant message from the previous round back into messages, including fields such as reasoning_content and tool_calls when they are present. This keeps the next request faithful to what the model actually produced.

Add streaming for better UX

For a web app or terminal assistant, waiting for the entire response can feel slow. The API supports streaming with stream: true in the JSON body.

import requests

url = "https://api.acedata.cloud/kimi/chat/completions"
headers = {
    "accept": "application/json",
    "authorization": "Bearer {token}",
    "content-type": "application/json"
}
payload = {
    "model": "kimi-k3",
    "messages": [{"role": "user", "content": "Hello"}],
    "reasoning_effort": "max",
    "stream": True
}

response = requests.post(url, json=payload, headers=headers)
print(response.text)
Enter fullscreen mode Exit fullscreen mode

The streaming response arrives as multiple data: blocks. During the stream, new content appears inside choices[].delta. K3 may stream reasoning_content as well as final content. The stream is complete when the data value is [DONE].

In a UI, treat these chunks as events: append delta.content to the visible answer, optionally handle delta.reasoning_content separately, and stop reading when you receive [DONE].

Keep multi-turn chat simple

You do not need a special session object for a basic multi-turn chat. Send previous turns in the messages array:

{
  "model": "kimi-k3",
  "messages": [
    {"role": "assistant", "content": "Hello! How can I help you today?"},
    {"role": "user", "content": "What model are you?"}
  ],
  "reasoning_effort": "max"
}
Enter fullscreen mode Exit fullscreen mode

The response shape remains the same: inspect choices, read the assistant message, and track usage. As your conversation gets longer, this is also where token usage becomes important. Logging usage.total_tokens early will save you debugging time later.

Handle errors before users see them

The guide documents several error categories worth mapping into clear application messages:

  • 400 token_mismatched: bad request, possibly missing or invalid parameters.
  • 400 api_not_implemented: bad request, possibly missing or invalid parameters.
  • 401 invalid_token: invalid or missing authorization token.
  • 429 too_many_requests: rate limit exceeded.
  • 500 api_error: server-side failure.

A typical error response includes success: false, an error object with code and message, and a trace_id. Log the trace_id; it is the field you will want when investigating a failed request.

A good first build

If I were adding Kimi K3 to an app, I would start with one non-streaming endpoint, log id and usage, then add streaming only after the basic response parser is stable. After that, I would add multi-turn history and make sure the full previous assistant message is preserved.

That order keeps the integration boring: request shape first, response parsing second, streaming third, conversation memory last.

For the original field reference and examples, see the Ace Data Cloud Kimi Chat Completion API guide: https://platform.acedata.cloud/documents/kimi-chat-completion-integration

Top comments (0)