Moonshot AI
The Chinese lab behind the open-weight Kimi models, including Kimi K3, a 1M-context reasoning model priced flat across its full context.
What it actually does
Moonshot's platform serves the Kimi model line through an API. Kimi K3, launched July 2026, is a 1M-context reasoning model at $3 input / $15 output per 1M tokens, priced flat across the full context with no long-context surcharge, unlike xAI and Gemini. Cheaper K2-line models remain available for simpler tasks.
There is no ongoing token free tier: you recharge at least $1 to start and get a $5 voucher after your first $5 in recharges, and file-extraction endpoints are temporarily free. Cache-hit input drops to $0.30 per 1M tokens on K3, a 90% discount for repeated prompts.
Two caveats shape real bills. K3 always reasons at maximum effort with no cheaper non-thinking mode, so the $15 output rate dominates costs, and the older Moonshot V1 model line sunsets August 31, 2026. As a China-based company, it raises the same data-residency questions as DeepSeek.
Moonshot AI vs its main rivals
The tools people actually weigh against Moonshot AI: DeepSeek, Alibaba Qwen, Zhipu GLM. Same criteria for every column, including where Moonshot AI loses.
| ZZhipu GLM | ||||
|---|---|---|---|---|
Pricing & access | ||||
| Free tier | No$5 voucher | NoNo (prepay) | No90-day trial | YesFree GLM tier |
| Starting price | $0.95 /1M in | $0.14 /1M in | $0.10 /1M in | ~$0.11 /1M in |
Models & capability | ||||
| Flagship reasoning price | $3/$15 | $0.44/$0.87 | $1.25/$3.75 | ~$0.60/$2.20 |
| Near-frontier reasoning | YesYes | YesYes | YesYes | YesYes |
| Cheaper non-thinking mode | NoK3 always reasons | Yesflash tier | YesFlash | YesAir/Flash |
Scaling & limits | ||||
| Max context window | 1M | 1M | 256K-1M | 200K |
| Flat pricing across full context | YesYes | YesYes | NoLength tiers | UnknownUnknown |
| Cache-hit discount | Yes90% off | Yes~99% off | YesYes | UnknownUnknown |
Developer & API | ||||
| REST API access | Paid | Paid | Paid | Free |
Data & compliance | ||||
| Open weights / self-host option | YesYes | YesYes | YesYes | YesYes |
| Non-China data processing | NoChina-hosted | NoChina-hosted | YesSingapore endpoint | NoChina-hosted |
- Kimi K3 is priced flat across the full 1M-token context with no long-context surcharge
- Cache-hit input drops to $0.30 per 1M tokens on K3, a 90% discount for repeated prompts
- Open weights let you move to self-hosting or a cheaper inference host without prompt rewrites
- K3 always reasons at max effort with no cheaper non-thinking mode, so the $15 output rate dominates real bills
- The older Moonshot V1 model line sunsets August 31, 2026, breaking code pinned to those names
- China-based hosting raises the same data-residency concerns as DeepSeek for US and EU customers
Pay less for it
5 ways foundK3 cache hits drop input to $0.30 per 1M tokens, a 90% cut on repeated prompts
Get a $5 voucher after your first $5 in recharges
Use the cheaper older Kimi models for simple tasks that do not need K3 reasoning
Run Kimi on a cheaper host or your own hardware to escape per-token pricing
DeepSeek — Similar China-hosted near-frontier reasoning at a fraction of K3's per-token cost, with a cheaper flash tier for simple tasks
StackTracker tracks what you actually pay for Moonshot AI and every other tool, flags overpayment, and shows the dollars you would save by switching.
Is it the right tool for you
- You want a frontier-class open-weight reasoning model with a self-host escape hatch
- Your workloads are long-context and you want flat pricing with no surcharge
- You can accept China-hosted inference for your customer base
- You run many simple tasks and want a cheap non-thinking mode → use DeepSeek or Alibaba Qwen
- Your contracts forbid China-based data processing → use Mistral or OpenAI
Track what Moonshot AI and the rest of your stack cost
StackTracker adds up every subscription, plus your hours, so you see the real number.
Prices and limits last verified 2026-07-20.