Moonshot AI logo

Moonshot AI

The Chinese lab behind the open-weight Kimi models, including Kimi K3, a 1M-context reasoning model priced flat across its full context.

Roughly $0.95-$3 /1M input depending on model (K2-line to Kimi K3), output up to $15 on K3

What it actually does

Moonshot's platform serves the Kimi model line through an API. Kimi K3, launched July 2026, is a 1M-context reasoning model at $3 input / $15 output per 1M tokens, priced flat across the full context with no long-context surcharge, unlike xAI and Gemini. Cheaper K2-line models remain available for simpler tasks.

There is no ongoing token free tier: you recharge at least $1 to start and get a $5 voucher after your first $5 in recharges, and file-extraction endpoints are temporarily free. Cache-hit input drops to $0.30 per 1M tokens on K3, a 90% discount for repeated prompts.

Two caveats shape real bills. K3 always reasons at maximum effort with no cheaper non-thinking mode, so the $15 output rate dominates costs, and the older Moonshot V1 model line sunsets August 31, 2026. As a China-based company, it raises the same data-residency questions as DeepSeek.

Moonshot AI vs its main rivals

The tools people actually weigh against Moonshot AI: DeepSeek, Alibaba Qwen, Zhipu GLM. Same criteria for every column, including where Moonshot AI loses.

Moonshot AI logoMoonshot AIDeepSeek logoDeepSeekAlibaba Qwen logoAlibaba QwenZZhipu GLM
Pricing & access
Free tierNo$5 voucherNoNo (prepay)No90-day trialYesFree GLM tier
Starting price$0.95 /1M in$0.14 /1M in$0.10 /1M in~$0.11 /1M in
Models & capability
Flagship reasoning price$3/$15$0.44/$0.87$1.25/$3.75~$0.60/$2.20
Near-frontier reasoningYesYesYesYesYesYesYesYes
Cheaper non-thinking modeNoK3 always reasonsYesflash tierYesFlashYesAir/Flash
Scaling & limits
Max context window1M1M256K-1M200K
Flat pricing across full contextYesYesYesYesNoLength tiersUnknownUnknown
Cache-hit discountYes90% offYes~99% offYesYesUnknownUnknown
Developer & API
REST API accessPaidPaidPaidFree
Data & compliance
Open weights / self-host optionYesYesYesYesYesYesYesYes
Non-China data processingNoChina-hostedNoChina-hostedYesSingapore endpointNoChina-hosted
Worth it for
  • Kimi K3 is priced flat across the full 1M-token context with no long-context surcharge
  • Cache-hit input drops to $0.30 per 1M tokens on K3, a 90% discount for repeated prompts
  • Open weights let you move to self-hosting or a cheaper inference host without prompt rewrites
Watch out for
  • K3 always reasons at max effort with no cheaper non-thinking mode, so the $15 output rate dominates real bills
  • The older Moonshot V1 model line sunsets August 31, 2026, breaking code pinned to those names
  • China-based hosting raises the same data-residency concerns as DeepSeek for US and EU customers

Pay less for it

5 ways found
Cache-hit discount

K3 cache hits drop input to $0.30 per 1M tokens, a 90% cut on repeated prompts

Recharge voucher

Get a $5 voucher after your first $5 in recharges

Route to K2-line

Use the cheaper older Kimi models for simple tasks that do not need K3 reasoning

Self-host open weights

Run Kimi on a cheaper host or your own hardware to escape per-token pricing

Cheaper swap

DeepSeekSimilar China-hosted near-frontier reasoning at a fraction of K3's per-token cost, with a cheaper flash tier for simple tasks

See this for your whole stack, with your real numbers

StackTracker tracks what you actually pay for Moonshot AI and every other tool, flags overpayment, and shows the dollars you would save by switching.

See plans

Is it the right tool for you

Pick it if
  • You want a frontier-class open-weight reasoning model with a self-host escape hatch
  • Your workloads are long-context and you want flat pricing with no surcharge
  • You can accept China-hosted inference for your customer base
Skip it if
  • You run many simple tasks and want a cheap non-thinking mode → use DeepSeek or Alibaba Qwen
  • Your contracts forbid China-based data processing → use Mistral or OpenAI

Track what Moonshot AI and the rest of your stack cost

StackTracker adds up every subscription, plus your hours, so you see the real number.

Prices and limits last verified 2026-07-20.