Best AI APIs & LLM providers for solo founders
If your product calls an LLM, token spend is usually your most volatile cost line: it scales with usage, not with plan tiers, and one chatty feature can double your bill overnight. Every provider here is pay-per-token with no seats and no contracts, and each one ships a ladder of models from cheap workhorses to frontier flagships, so compare price ranges per provider rather than a single model. For judging model quality per use case, skip the marketing pages and check neutral leaderboards like LMArena (lmarena.ai) and Artificial Analysis, then match the cheapest model that clears your quality bar.
| Tool | Free tier | Paid from | Best for |
|---|---|---|---|
| $0.20-$5 /1M input | Founders who want the safest default with the most tutorials and a cheap nano tier for high-volume tasks | ||
| $1-$10 /1M input | Founders building coding tools or agents who will actually implement prompt caching | ||
| $0.10-$4 /1M input | Pre-revenue founders who want to validate an AI feature for literally $0 | ||
| $1-$4 /1M input | Founders who want frontier-class output below Claude and OpenAI flagship prices and are fine with the data-sharing tradeoff for credits | ||
| $0.14-$0.44 /1M input | Cost-obsessed founders whose users will not object to a China-hosted model | ||
| roughly $0.95-$3 /1M input | Founders who want a frontier-class open-weight reasoning model and long-context workloads without surcharges | ||
| $0.10-$2.50 /1M input | Founders who standardized on Qwen models and want first-party pricing with a self-host option later | ||
| $0.15-$1.50 /1M input | EU-based founders or anyone who wants a self-host escape hatch | ||
| $0.05-$1 /1M input | Founders whose feature works on open models and who care about speed and cost above all | ||
Together AI$0.17-$1.74 /1M input | $0.17-$1.74 /1M input | Founders who need many open models plus fine-tuning under one roof and expect to scale to dedicated capacity | |
| pass-through of each model's provider price across 400+ models | Founders still choosing a model, or apps that let users pick their own model |
Judge quality with neutral leaderboards, not vendor pages: check LMArena (lmarena.ai) and Artificial Analysis for your specific use case, then buy the cheapest model that clears the bar. Start on Gemini's free tier or Groq's no-card free tier to validate the feature for $0, then put your production default on a cheap workhorse: Gemini 2.5 Flash-Lite ($0.10/$0.40), Qwen3.5 Flash ($0.10/$0.40), gpt-5.4-nano ($0.20/$1.25), or DeepSeek v4-flash ($0.14/$0.28) per 1M tokens.
For coding or agent features, pay up for Claude Sonnet 5 at the $2/$10 intro price but budget for $3/$15 after August 2026, and consider Grok 4.5 at $2/$6 as the value flagship. Turn on prompt caching before you scale, it cuts repeated input to 10-25% on most providers.
A typical solo SaaS with a few thousand AI interactions a month lands at $5-30/mo on a cheap model or $30-150/mo on a frontier model. Use OpenRouter only while comparing; once you commit, go direct and skip its 5.5% credit fee.
Ask about these tools
The fine print, tool by tool
OpenAI$0.20-$5 /1M input+
The default LLM API most tutorials assume, with GPT models from a cheap nano tier up to frontier reasoning models. Solo founders pick it when they want the largest ecosystem of examples and SDKs and a wide price ladder inside one account.
$0.20-$5 /1M input depending on model (gpt-5.4-nano to gpt-5.6-sol), output $1.25-$30; pro models up to $30 input / $180 output
No free tier. Priority processing costs 2-2.5x standard rates. Cache writes are charged (20-25% of output rates). Data residency adds 10% on models released after March 2026. Batch discount only on select models.
- + Widest model ladder for cost routing: gpt-5.4-nano at $0.20/$1.25 up to gpt-5.6-sol at $5/$30 per 1M tokens
- + Prompt caching cuts cached input to 10% of the standard rate, huge for repeated system prompts
- + Batch API and Flex processing give 50% off input and output for non-urgent jobs
- − No API free tier at all, unlike Gemini and Groq, so even testing costs money
- − Pro-grade models are a trap for solo budgets: gpt-5.5-pro and gpt-5.4-pro run $30 input / $180 output per 1M tokens
Anthropic Claude$1-$10 /1M input+
Anthropic's API for the Claude family, which leads most coding and agent benchmarks. Solo founders pick it when the product is a coding tool or agent workflow and they are willing to pay above commodity rates for output quality.
$1-$10 /1M input depending on model (Haiku 4.5 to Fable 5 and Opus fast mode), output $5-$50
Sonnet 5 $2/$10 is introductory only, becomes $3/$15 after Aug 31, 2026. 1-hour cache writes cost 2x input price. US-only inference (inference_geo) adds a 1.1x multiplier. Web search costs $10 per 1,000 searches. Opus 4.8 fast mode doubles prices to $10/$50.
- + Sonnet 5 intro pricing of $2/$10 per 1M tokens through August 31, 2026 is strong value for a frontier model
- + Cache reads cost 10% of input price and batch API takes another 50% off, and the discounts stack
- + Best-in-class at coding and agent workloads, which is what most solo founders actually build
- − Cheapest model (Haiku 4.5 at $1/$5) costs 5x OpenAI's nano tier and 10x Gemini Flash-Lite for simple tasks
- − Sonnet 5 jumps 50% to $3/$15 per 1M tokens on September 1, 2026, so budget for the real price
Google Gemini$0.10-$4 /1M input+
Google's LLM API with the only real ongoing free tier among the big labs, plus some of the cheapest paid models anywhere. Solo founders pick it to validate an AI feature at zero cost and often stay for Flash-Lite pricing in production.
$0.10-$4 /1M input depending on model (2.5 Flash-Lite to 3.1 Pro at long-prompt rates), output $0.40 and up
genuinely free tier with no card: free daily usage of Flash models and embeddings, rate-limited per model
Free tier data is used for product improvement; paid tier data is not. Tier 1 caps monthly spend at $250 until you have paid $100+ (plus a 3-day wait). Pro model input price roughly doubles above the long-prompt threshold (Gemini 3.1 Pro goes from $2 to $4 per 1M input). Priority inference costs 1.8x standard.
- + Only major frontier lab with a real ongoing free API tier, enough to build and demo an MVP at $0
- + Gemini 2.5 Flash-Lite at $0.10/$0.40 per 1M tokens is the cheapest big-lab model in this list
- + Batch API gives a flat 50% discount and context caching cuts repeated-input costs
- − Free tier prompts can be used to train Google's models, a real issue if users paste private data
- − Free tier blocks context caching, batch processing, and search grounding, so your cost optimizations only exist on paid
xAI$1-$4 /1M input+
The API for xAI's Grok models, from the cheap Grok 4.1 Fast to the Grok 4.5 flagship released July 2026. Solo founders pick it for competitive frontier pricing ($2/$6 on Grok 4.5) and the monthly credit program, if they accept sending prompt data to xAI to earn it.
$1-$4 /1M input depending on model (grok-build-0.1 to Grok 4.5 at long-context rates), output $2-$12
Grok 4.3 sits at $1.25/$2.50 per 1M tokens as the mid-tier. Models with two price rows use long-context pricing: crossing the token threshold reprices all tokens in the request. Data-sharing credits require $5 prior spend and are limited to eligible countries.
- + grok-build-0.1 at $1/$2 per 1M tokens is the cheapest way into the Grok line for high-volume tasks
- + Grok 4.5 at $2/$6 per 1M tokens undercuts most frontier flagships, with cached input at $0.30 (85% off)
- + Data-sharing program pays up to $150/mo in API credits, enough to run a small production feature free
- − Long-context billing is punitive: prompts at or above 200K tokens bill the whole request at the higher rate, $4/$12 on Grok 4.5
- − The free credits require opting in to data sharing, so your users' prompts help train xAI models
DeepSeek$0.14-$0.44 /1M input+
A China-based lab selling near-frontier reasoning models at commodity prices through an OpenAI-compatible API. Solo founders pick it when cost per token matters more than anything and their customers have no objection to China-hosted inference.
$0.14-$0.44 /1M input (cache miss) depending on model (v4-flash to v4-pro), output $0.28-$0.87
Billing counts cache-hit vs cache-miss input separately, so real cost depends heavily on prompt reuse. Concurrency capped at 2,500 requests (flash) and 500 (pro). Older off-peak discounts no longer appear on the pricing page. Watch the July 2026 model-name deprecation.
- + Near-frontier reasoning at commodity prices: v4-pro is $0.435 input / $0.87 output per 1M tokens, roughly 10x cheaper than Western flagships
- + Automatic cache hits drop input to $0.0028 per 1M tokens on flash, effectively free for repeated prompts
- + 1M-token context with up to 384K output on both models at no surcharge
- − China-based data processing is a hard blocker for many US and EU customer contracts, check before you build on it
- − Model names deepseek-chat and deepseek-reasoner are deprecated as of July 24, 2026, code pinned to them breaks
Moonshot AIroughly $0.95-$3 /1M input+
The Chinese lab behind the open-weight Kimi models, including Kimi K3, a 1M-context reasoning model launched July 2026 that competes with Western flagships. Solo founders pick it for frontier-adjacent reasoning on an open-weight model with a self-host escape hatch.
roughly $0.95-$3 /1M input depending on model (K2-line to Kimi K3), output up to $15 on K3
Kimi K3 is $3 input / $15 output per 1M tokens, cache hits $0.30. China-based company; weigh the same data-residency questions as DeepSeek for US and EU customers. Cheaper K2-line models remain available for simple tasks.
- + Kimi K3 pricing is flat across the full 1M-token context with no long-context surcharge, unlike xAI and Gemini
- + Cache-hit input drops to $0.30 per 1M tokens on K3, a 90% discount for repeated prompts
- + Open weights mean you can move to self-hosting or a cheaper inference host without prompt rewrites
- − K3 always reasons at max effort with no cheaper non-thinking mode, so the $15/1M output rate dominates real bills
- − The older Moonshot V1 model line sunsets August 31, 2026, and code pinned to those names breaks
Alibaba Cloud (Qwen)$0.10-$2.50 /1M input+
Alibaba's Model Studio serves the Qwen family, the most widely used open-weight model line, through an OpenAI-compatible API. Solo founders pick it when they want Qwen quality at first-party prices, especially the $0.05 Flash tier for high-volume simple tasks.
$0.10-$2.50 /1M input depending on model (Qwen3.5 Flash to Qwen3.7-Max list price), output $0.40-$7.50
Use the international (Singapore) endpoint, not the China mainland one, for products serving Western users. The Qwen3.7-Max $1.25/$3.75 rate is a 50% promotional discount off a $2.50/$7.50 list price, so budget for list. Alibaba Cloud account required, which adds signup friction versus a simple API key.
- + Qwen3.5 Flash at $0.10 input / $0.40 output per 1M tokens is among the cheapest usable models anywhere
- + Batch calls cut both input and output to 50% of real-time prices on supported models
- + Qwen3.7-Max flagship runs $1.25/$3.75 per 1M tokens on a 50% promo, well under Western flagship rates
- − The ongoing developer free tier ended April 15, 2026, leaving only a 90-day 1M-token-per-model trial
- − Length-based pricing bites: Qwen3-Max jumps from $1.20/$6.00 to $3.00/$15.00 per 1M tokens on prompts past 128K
Mistral$0.15-$1.50 /1M input+
A French lab selling open-weight models with EU-hosted inference through La Plateforme. Solo founders pick it when GDPR and EU data residency questions matter to their customers, or when they want the option to self-host the exact same weights later.
$0.15-$1.50 /1M input depending on model (Small 4 to Medium 3.5), output $0.60-$7.50
free experimentation tier on La Plateforme with strict fair-use rate limits; the generous Free plan is for Le Chat, not the API
Batch is 50% off, caching is 90% off input. OCR is priced per page ($4 per 1,000 pages), not per token. Free-tier usage is subject to fair-use limits defined only in the ToS. Student discount capped at 12 months.
- + Mistral Large 3 at $0.50/$1.50 per 1M tokens is very cheap for a large open-weight multimodal model
- + Input caching gives a 90% discount and batch processing halves the price on top
- + EU-hosted servers and open weights: simple GDPR answers now, self-host escape hatch later
- − Flagship Medium 3.5 at $1.50/$7.50 per 1M tokens trails GPT and Claude flagships on hard reasoning, so you pay near-frontier prices for sub-frontier output
- − Smaller ecosystem: fewer SDK integrations, templates, and community fixes than OpenAI or Anthropic
Groq$0.05-$1 /1M input+
An inference platform that runs open-weight models (Llama, GPT OSS, Qwen, and others) on its own LPU chips at very high speed, and a different company from xAI's Grok models despite the similar name. Solo founders pick it when their feature works on open models and response speed is part of the product.
$0.05-$1 /1M input depending on model (Llama 3.1 8B to larger open models), output from $0.08
free tier with no credit card: roughly 30 requests/min, 6,000 tokens/min, and about 1,000 requests/day on most models
Prompt caching gives 50% off cached input, batch API another 50% off with a 24-hour to 7-day window. Speech pricing is per hour of audio ($0.04-$0.111/hr), TTS is per million characters ($22+). OpenAI-compatible API, and adding a card unlocks roughly 10x rate limits with no minimum spend.
- + Cheapest tokens in this comparison: Llama 3.1 8B at $0.05/$0.08 per 1M tokens
- + Extreme speed (up to 1,000 tokens/sec on GPT OSS 20B) makes chat UIs feel instant
- + No-card free tier is enough to ship a small production feature, not just a demo
- − Open-weight models only, no GPT, Claude, or Gemini, so quality tops out below frontier for hard reasoning
- − Free tier throttles hard: 6,000 tokens/min means one long-context request can eat a whole minute
Together AI$0.17-$1.74 /1M input+
An inference platform hosting 200+ open models behind one API, plus fine-tuning and dedicated GPUs when you outgrow serverless. Solo founders pick it when they need many open models and fine-tuning in one place rather than the lowest per-token price.
$0.17-$1.74 /1M input depending on model (Qwen 3.5 9B to DeepSeek V4 Pro), output $0.25-$3.48
$25 signup credit was retired in 2025; today there are no free credits and a $5 minimum top-up. Serverless per-token prices are often 1.5-4x cheaper elsewhere for identical models (DeepSeek V4 Pro is 4-12x DeepSeek direct), the value is the platform. Provisioned throughput bills per PTU-minute even when idle.
- + 200+ open models behind one API, including DeepSeek, Qwen, Kimi, and Llama the day they drop
- + Fine-tuning from $0.48 per 1M training tokens lets you own a small custom model cheaply
- + Clear upgrade path from serverless tokens to dedicated endpoints and H100 GPUs at $3.99/hr when you scale
- − Per-token rates run above rivals for the same open models: Llama 3.3 70B is $1.04/$1.04 here vs $0.59/$0.79 on Groq
- − No real free tier and a $5 minimum purchase, worse for testing than Groq or Gemini
OpenRouterpass-through of each model's provider price across 400+ models+
A router that puts 400+ models from 70+ providers behind one API key at pass-through prices, funded by a fee on credit top-ups. Solo founders pick it while still comparing models, or when the app lets users choose their own model.
pass-through of each model's provider price across 400+ models, plus about 5.5% on credit purchases
free models limited to 50 requests/day without credits, 1,000 requests/day once you have bought $10+ in credits
The 5.5% Stripe fee (5% crypto) applies to credit top-ups, not per request, so top up in larger chunks. BYOK costs 5% of the normal model price after the first 1M requests/month. Once you settle on one model long-term, going direct saves the fee.
- + One API key for 400+ models from 70+ providers, swap GPT for Claude for DeepSeek by changing a string
- + No markup on inference itself, model prices are passed through from providers
- + Free-model catalog is a zero-cost way to benchmark which model your feature actually needs
- − 5.5% fee ($0.80 minimum) on every Stripe credit purchase is a pure tax versus going direct
- − Free models cap at 50 requests/day across all free models combined, useless for production
[ Partnerships ] Building a tool solo founders should know about?
DM @monjodavKnow what your whole stack costs
StackTracker adds up every one of these, plus your hours.
Start tracking · $9/mo