Kimi K3 for Coding: Pricing, Benchmarks, and Tool Setup
Review Kimi K3 API pricing, coding benchmark caveats, Token usage signals, and setup paths for Claude Code, Codex CLI, and OpenCode.
Contents
Kimi K3 is worth testing for long agent-based tasks, working with a large repository, and scenarios where images are added to the code. But the model should not be chosen from one place in the table: its results depend on harness, maximum reasoning, endpoint and the task itself, and constant thinking can make a short cycle slow and expensive.
Check the current model price before testing.
| If your script | What to check with Kimi K3 | When to choose another model |
|---|---|---|
| Long agent task in a large repository | Does the model hold the task boundaries and complete the check | If an agent goes into unnecessary actions or requires constant intervention |
| Short, frequent edits | Response time and reasoning Token cost | If a lighter model gives the same result faster and at lower cost |
| Working with code and images | Whether native vision helps with your input data | If the image does not affect the acceptance criteria |
What Moonshot AI released
According to the official repository, Kimi K3 is a sparse MoE model with 2.8 trillion parameters, of which 104 billion are activated on Token. Context - 1,048,576 Tokens. The model accepts text, images, and video, and reasoning is always enabled; in the published benchmark configuration, maximum effort was used.
These numbers describe the model’s capacity and mode, but do not guarantee that every token of a large codebase will be useful. As context grows, test whether the agent finds the right files, respects task boundaries, and completes a long tool chain.
How much does one run cost?
Rates and availability change, so check the selected provider’s catalog on the day of the test. To estimate one run, multiply actual input and output Tokens by the corresponding current rates. If a cache-hit rate is listed separately, apply it only to the confirmed cached volume; do not transfer a discount between providers.
Use current rates for your calculation, not an old comparison. Open current BetterToken prices
With both the official API and a compatible provider, total cost is not determined by input alone: long reasoning is billed as output and can become the main expense.
What coding benchmarks show
The official Moonshot table published, among other things, the following results:
- Terminal-Bench 2.1: 88.3.
- DeepSWE: 67.5.
- ProgramBench: 77.8.
- FrontierSWE: 81.2.
- SWE-Marathon: 42.0.
- Kimi Code Bench 2.0: 72.9.
These are vendor-reported results. For Kimi, the Kimi Code harness was often used, and for some competitors, their best published harness was used. The footnotes also indicate the maximum effort and fixed generation parameters. Therefore, a difference of several points does not answer the question of which model is better with your CLI, set of tools and project rules.
A useful unit of comparison looks like this:
model + harness + effort + endpoint + task.
If at least one element changes, the previous score becomes only a guide. An independent test is more useful when it captures the same repository, the same time limit, the same tools and the same acceptance criteria.
For what tasks does Kimi K3 look appropriate?
A strong candidate is a task where the model must hold the architecture for a long time, move between files and check the result itself. For example:
- parse a large repository and create a dependency map;
- fix a bug that goes through several modules;
- compare the layout or screenshot with the implementation of the interface;
- carry out a long plan with tests and re-checking;
- explore several options before changing the code.
A short interactive cycle looks weaker: rename the function, fix one check, explain a small diff. Always-on thinking adds delay and Token output even where long reasoning is not needed. For such problems, first compare K3 with an easier model using the same five examples.
Connect to Claude Code
Moonshot provides an Anthropic-compatible endpoint. The minimum configuration of the current international platform looks like this:
export ANTHROPIC_BASE_URL="https://api.moonshot.ai/anthropic"
export ANTHROPIC_AUTH_TOKEN="YOUR_API_KEY"
export ANTHROPIC_MODEL="kimi-k3[1m]"
export ANTHROPIC_DEFAULT_HAIKU_MODEL="kimi-k3[1m]"
export CLAUDE_CODE_SUBAGENT_MODEL="kimi-k3[1m]"
Once running, run /status. The response must have the same Base URL and Model ID. Checking through the /model menu is not sufficient: the interface may show Claude Code’s own aliases rather than the actual model of the compatible endpoint.
Do not mix Moonshot Open Platform and Kimi Code keys with someone else’s Base URL. Platforms have different domains and key types; 401 in this configuration does not indicate the quality of the model.
Connecting to Codex CLI
Moonshot supports the Responses API, so Codex connects directly to Kimi K3 without a local proxy or protocol conversion. Store the key in KIMI_API_KEY, not in config.toml, then add this to ~/.codex/config.toml:
model = "kimi-k3"
model_provider = "kimi"
model_context_window = 1048576
[model_providers.kimi]
name = "Kimi"
base_url = "https://api.moonshot.ai/v1"
env_key = "KIMI_API_KEY"
wire_api = "responses"
Restart Codex completely, confirm that the current model is kimi-k3, then send a short request without changing files. Codex Desktop and CLI read the same user config. Codex has no built-in video channel: for video, use the direct Kimi API as documented rather than treating manually extracted frames as the original input.
Connect to OpenCode
The official way is shorter:
opencode auth login
Select Moonshot AI, insert the key via secure dialog, then use /models to select Kimi K3 and /variants to check effort. The test should force the model to read a couple of files and return their names, but not change the project. After answering, check the selected model and Token consumption in the provider’s account.
What problems have already been described by users
In issue MoonshotAI/kimi-code #1911, the user described freezing, poor response to stopping, and going beyond the task boundaries: instead of reporting, the agent started a long chain of shell commands. In #2031 the author reported a session with 18.3 million input Tokens.
These are two separate user posts and not an independent reproduction or statistics of the entire model. It’s useful to turn them into checks for your own run:
- set explicit file boundaries and prohibit collateral changes;
- limit the number of iterations and subagent;
- stop the launch if several steps do not bring you closer to the readiness criterion;
- measure input, output, reasoning and time, and not just price per million;
- start from a separate branch or test repository.
How to conduct your own short test
- Prepare five tasks from your work: one local bug, one change across several files, one test, one repository search, and one image-based task.
- Use the same rules, time limit, and PASS criterion for every model.
- Record total input tokens, output tokens, duration, manual interventions, and any boundary violations.
- After five runs, verify each result against the PASS criterion, compare token use with the baseline, and decide whether Kimi K3 fits as a primary agent, a model for complex tasks, or only a backup option.