초대하고 적립

초대 보상 안내

초대 링크를 공유하세요. 친구가 링크로 가입하고 충전하면 이후 충전마다 표시된 보상을 받을 수 있습니다.

Kimi K3 코딩 평가: 비용, benchmark, 도구 설정

Kimi K3를 API 비용, vendor benchmark, 도구 설정과 scope·latency·token 사용 검증 기준으로 평가하는 가이드입니다.

목차

Kimi K3는 긴 coding run, 대규모 repository, code와 visual input을 결합하는 task의 후보입니다. leaderboard만 보고 선택하는 것은 위험합니다. 결과는 harness, reasoning effort, endpoint, task에 따라 달라집니다. 항상 활성화된 thinking 때문에 짧은 loop가 예상보다 느리고 비싸질 수도 있습니다.

테스트 전에 현재 모델 가격을 확인하세요.

WorkloadKimi K3로 검증할 항목다른 모델을 선택할 때
대규모 repository의 긴 agent runscope를 지키고 verification을 완료하는지task를 확장하거나 반복적인 개입이 필요할 때
빈번한 작은 editresponse time과 reasoning-output cost가벼운 모델이 같은 결과를 더 빨리 낼 때
code와 visual inputnative vision이 실제 acceptance test를 개선하는지image가 pass condition에 영향을 주지 않을 때

Moonshot AI가 공개한 것

공식 Kimi K3 repository는 총 2.8 trillion parameter 중 token마다 104 billion이 활성화되는 sparse MoE 모델로 설명합니다. 1,048,576-token context window, native text, image, video input, 계속 활성화되는 thinking을 지원합니다.

큰 context window는 capacity이지 repository의 모든 token이 효과적으로 사용된다는 보장이 아닙니다. 유용한 평가는 agent가 올바른 file을 찾고, scope를 지키며, tool error 후 복구하고, budget 안에서 완료하는지 확인해야 합니다.

테스트 전에 API 비용 계산하기

가격과 availability는 바뀌므로 test 당일 선택한 provider의 catalog를 확인하세요. 한 run의 비용은 실제 input과 output token에 각각의 current rate를 곱해 추정합니다. cache-hit rate가 별도로 있으면 confirmed cached volume에만 적용하고 provider 사이에 discount를 옮기지 마세요.

계산에는 오래된 비교가 아니라 current rate를 사용하세요. BetterToken 현재 가격 열기

Official API와 compatible provider 모두 total cost는 input만으로 정해지지 않습니다. 긴 reasoning은 output으로 과금되며 가장 큰 비용이 될 수 있습니다.

coding benchmark는 harness와 함께 읽기

Moonshot이 보고한 Kimi K3 score에는 Terminal-Bench 2.1의 88.3, DeepSWE의 67.5, ProgramBench의 77.8, FrontierSWE의 81.2, SWE-Marathon의 42.0, Kimi Code Bench 2.0의 72.9가 포함됩니다. 이는 vendor-reported result입니다. repository에는 Kimi K3가 max reasoning effort를 사용하며, 일부 coding evaluation은 Kimi Code harness를 사용한 반면 competitor는 다른 published harness를 사용할 수 있다고 적혀 있습니다.

유용한 comparison unit은 다음과 같습니다.

model + harness + effort + endpoint + task.

score만으로 agent가 사용한 token 수, 요청한 file 안에 머물렀는지, 사람이 몇 번 방향을 바로잡아야 했는지는 알 수 없습니다.

Kimi K3를 coding tool에 연결하기

Moonshot은 Claude Code, Codex CLI, OpenCode를 위한 공식 guide를 각각 제공합니다. 하나의 configuration이 세 tool에서 변경 없이 작동한다고 가정하지 마세요.

  1. 사용할 tool과 공식 문서에 명시된 protocol 경로를 선택하세요. 다른 client의 endpoint를 재사용하지 마세요.
  2. tool의 credential mechanism을 통해 API key를 입력한 뒤, 아래 문서에 나온 endpoint와 model을 정확히 설정하세요.
  3. client를 재시작하고 read-only 작업 하나를 실행한 다음, 파일 변경을 허용하기 전에 실제 적용된 model, Base URL, usage를 확인하세요.

Claude Code의 문서화된 international Anthropic-compatible setup은 다음을 사용합니다.

export ANTHROPIC_BASE_URL="https://api.moonshot.ai/anthropic"
export ANTHROPIC_AUTH_TOKEN="YOUR_API_KEY"
export ANTHROPIC_MODEL="kimi-k3[1m]"
export ANTHROPIC_DEFAULT_HAIKU_MODEL="kimi-k3[1m]"
export CLAUDE_CODE_SUBAGENT_MODEL="kimi-k3[1m]"

client를 다시 시작하고 /status를 실행해 실제 Base URL과 model을 확인하세요. key를 commit하거나 issue에 붙여 넣지 마세요.

Moonshot Open Platform과 Kimi Code의 key를 다른 platform의 Base URL과 섞지 마세요. domain과 key type이 다르므로 이 configuration의 401은 model 품질을 뜻하지 않습니다.

Moonshot은 Responses API를 지원하므로 Codex는 local proxy나 protocol conversion 없이 Kimi K3에 직접 연결합니다. key는 config.toml이 아니라 KIMI_API_KEY에 저장하고 ~/.codex/config.toml에 다음을 추가하세요:

model = "kimi-k3"
model_provider = "kimi"
model_context_window = 1048576

[model_providers.kimi]
name = "Kimi"
base_url = "https://api.moonshot.ai/v1"
env_key = "KIMI_API_KEY"
wire_api = "responses"

Codex를 완전히 restart하고 current model이 kimi-k3인지 확인한 뒤 file을 바꾸지 않는 짧은 request를 보내세요. Desktop과 CLI는 같은 user config를 읽습니다. Codex에는 built-in video channel이 없으므로 video에는 documented direct Kimi API를 사용하세요.

OpenCode에서는 opencode auth login을 실행하고 Moonshot AI를 선택한 뒤 credential dialog에서 key를 입력합니다. 그런 다음 /models와 /variants로 Kimi K3와 effort를 선택하세요. read-only task로 시작하고 provider dashboard에서 model과 usage를 검증합니다.

사용자 report를 test case로 바꾸기

MoonshotAI/kimi-code issue #1911에서 한 user는 hang, 약한 interrupt, scope expansion을 설명합니다. #2031에서 작성자는 input Tokens 18.3 million인 session을 보고했습니다. 이는 별개의 두 user report이며 전체 model의 error rate나 independent reproduction이 아닙니다. 명시적인 file boundary, iteration limit, cancellation, token accounting, failed tool call 후 recovery를 테스트해야 하는 이유입니다.

대표 task 다섯 개를 준비하세요. local bug, multi-file change, test 추가, repository search, visual-input task 하나입니다. commit, prompt, tool, timeout, acceptance test를 고정합니다. input, cache, output, reasoning, duration, manual intervention, scope violation을 기록하세요. 결과를 보면 Kimi K3가 default agent, 어려운 task용 모델, fallback 중 어디에 적합한지 판단할 수 있습니다.

Sources

LLM 워크플로를 최적화할 준비가 되셨나요?

하나의 API로 모델을 연결하고 키와 AI 비용을 관리하세요.

무료로 시작하기