DeepSeek logo

DeepSeek

A China-based lab selling near-frontier reasoning models at commodity per-token prices through an OpenAI-compatible API.

$0.14-$0.44 /1M input (cache miss) depending on model (v4-flash to v4-pro), output $0.28-$0.87

What it actually does

DeepSeek offers an OpenAI-compatible API to its v4 models, from v4-flash at $0.14 input (cache miss) to v4-pro at $0.435/$0.87 per 1M tokens, roughly 10x cheaper than Western flagships for comparable reasoning quality. There is no free tier; you prepay credits, though the consumer chat app is free.

Billing counts cache-hit and cache-miss input separately, and automatic cache hits can drop input to $0.0028 per 1M tokens on flash, effectively free for repeated prompts. Both models offer a 1M-token context with up to 384K output at no surcharge.

The main blocker is data location: China-based processing rules out many US and EU customer contracts, so check before building on it. Also watch model-name deprecation, the deepseek-chat and deepseek-reasoner names are deprecated as of July 24, 2026.

DeepSeek vs its main rivals

The tools people actually weigh against DeepSeek: Alibaba Qwen, Moonshot Kimi, OpenAI. Same criteria for every column, including where DeepSeek loses.

DeepSeek logoDeepSeekAlibaba Qwen logoAlibaba QwenMoonshot Kimi logoMoonshot KimiOpenAI logoOpenAI
Pricing & access
Free tierNoNo (prepay)No90-day trialNoNo (voucher)NoNo
Starting price$0.14 /1M in$0.10 /1M in$0.95 /1M in$0.20 /1M in
Models & capability
Flagship reasoning price$0.44/$0.87$1.25/$3.75$3/$15up to $30/$180
Near-frontier reasoningYesYesYesYesYesYesYesFrontier
Scaling & limits
Max context window1M256K-1M1M400K
Automatic cache-hit discountYes~99% offYesYesYes90% offYes10% rate
Concurrency ceiling2,500 / 500UnknownUnknownHigh tiers
Developer & API
REST API accessPaidPaidPaidPaid
OpenAI-compatible endpointYesYesYesYesYesYesYesNative
Data & compliance
Open weights / self-host optionYesYesYesYesYesYesNoNo
Non-China data processingNoChina-hostedYesSingapore endpointNoChina-hostedYesUS/EU
Worth it for
  • Near-frontier reasoning at commodity prices: v4-pro at $0.435/$0.87 per 1M tokens, roughly 10x cheaper than Western flagships
  • Automatic cache hits drop input to $0.0028 per 1M tokens on flash, effectively free for repeated prompts
  • 1M-token context with up to 384K output on both models at no surcharge
Watch out for
  • China-based data processing is a hard blocker for many US and EU customer contracts
  • No free tier, so you prepay credits before any API testing
  • deepseek-chat and deepseek-reasoner model names are deprecated as of July 24, 2026

Pay less for it

4 ways found
Automatic cache hits

Repeated-prompt input drops to ~$0.0028 per 1M tokens on flash, near-free

Route to v4-flash

Use the $0.14/$0.28 flash tier for simple, high-volume tasks

Benchmark via OpenRouter free catalog

Test DeepSeek models at $0 through OpenRouter's free-model list before committing spend

Self-host open weights

Move to a cheaper inference host or your own GPUs without prompt rewrites

See this for your whole stack, with your real numbers

StackTracker tracks what you actually pay for DeepSeek and every other tool, flags overpayment, and shows the dollars you would save by switching.

See plans

Is it the right tool for you

Pick it if
  • Cost per token matters more than anything else in your build
  • Your customers have no objection to China-hosted inference
  • You want near-frontier reasoning with heavy prompt reuse to exploit cache hits
Skip it if
  • Your contracts forbid China-based data processing → use OpenAI, Mistral, or Alibaba Qwen's Singapore endpoint
  • You need a frontier coding or agent model → use Anthropic Claude

Track what DeepSeek and the rest of your stack cost

StackTracker adds up every subscription, plus your hours, so you see the real number.

Prices and limits last verified 2026-07-20.