DeepSeek
A China-based lab selling near-frontier reasoning models at commodity per-token prices through an OpenAI-compatible API.
What it actually does
DeepSeek offers an OpenAI-compatible API to its v4 models, from v4-flash at $0.14 input (cache miss) to v4-pro at $0.435/$0.87 per 1M tokens, roughly 10x cheaper than Western flagships for comparable reasoning quality. There is no free tier; you prepay credits, though the consumer chat app is free.
Billing counts cache-hit and cache-miss input separately, and automatic cache hits can drop input to $0.0028 per 1M tokens on flash, effectively free for repeated prompts. Both models offer a 1M-token context with up to 384K output at no surcharge.
The main blocker is data location: China-based processing rules out many US and EU customer contracts, so check before building on it. Also watch model-name deprecation, the deepseek-chat and deepseek-reasoner names are deprecated as of July 24, 2026.
DeepSeek vs its main rivals
The tools people actually weigh against DeepSeek: Alibaba Qwen, Moonshot Kimi, OpenAI. Same criteria for every column, including where DeepSeek loses.
Pricing & access | ||||
| Free tier | NoNo (prepay) | No90-day trial | NoNo (voucher) | NoNo |
| Starting price | $0.14 /1M in | $0.10 /1M in | $0.95 /1M in | $0.20 /1M in |
Models & capability | ||||
| Flagship reasoning price | $0.44/$0.87 | $1.25/$3.75 | $3/$15 | up to $30/$180 |
| Near-frontier reasoning | YesYes | YesYes | YesYes | YesFrontier |
Scaling & limits | ||||
| Max context window | 1M | 256K-1M | 1M | 400K |
| Automatic cache-hit discount | Yes~99% off | YesYes | Yes90% off | Yes10% rate |
| Concurrency ceiling | 2,500 / 500 | Unknown | Unknown | High tiers |
Developer & API | ||||
| REST API access | Paid | Paid | Paid | Paid |
| OpenAI-compatible endpoint | YesYes | YesYes | YesYes | YesNative |
Data & compliance | ||||
| Open weights / self-host option | YesYes | YesYes | YesYes | NoNo |
| Non-China data processing | NoChina-hosted | YesSingapore endpoint | NoChina-hosted | YesUS/EU |
- Near-frontier reasoning at commodity prices: v4-pro at $0.435/$0.87 per 1M tokens, roughly 10x cheaper than Western flagships
- Automatic cache hits drop input to $0.0028 per 1M tokens on flash, effectively free for repeated prompts
- 1M-token context with up to 384K output on both models at no surcharge
- China-based data processing is a hard blocker for many US and EU customer contracts
- No free tier, so you prepay credits before any API testing
- deepseek-chat and deepseek-reasoner model names are deprecated as of July 24, 2026
Pay less for it
4 ways foundRepeated-prompt input drops to ~$0.0028 per 1M tokens on flash, near-free
Use the $0.14/$0.28 flash tier for simple, high-volume tasks
Test DeepSeek models at $0 through OpenRouter's free-model list before committing spend
Move to a cheaper inference host or your own GPUs without prompt rewrites
StackTracker tracks what you actually pay for DeepSeek and every other tool, flags overpayment, and shows the dollars you would save by switching.
Is it the right tool for you
- Cost per token matters more than anything else in your build
- Your customers have no objection to China-hosted inference
- You want near-frontier reasoning with heavy prompt reuse to exploit cache hits
- Your contracts forbid China-based data processing → use OpenAI, Mistral, or Alibaba Qwen's Singapore endpoint
- You need a frontier coding or agent model → use Anthropic Claude
Track what DeepSeek and the rest of your stack cost
StackTracker adds up every subscription, plus your hours, so you see the real number.
Prices and limits last verified 2026-07-20.