Groq
An inference platform running open-weight models on custom LPU chips optimized for token throughput, distinct from xAI's Grok models.
What it actually does
Groq serves open-weight models (Llama, GPT OSS, Qwen, and others) on its own LPU hardware, optimized for token throughput. Prices start at $0.05/$0.08 per 1M tokens on Llama 3.1 8B, the cheapest tokens in this comparison, and it is a different company from xAI despite the similar name.
It has a no-credit-card free tier of roughly 30 requests/min, 6,000 tokens/min, and about 14,400 requests/day on most models, enough to ship a small production feature rather than just a demo. Adding a card unlocks roughly 10x rate limits with no minimum spend, and the API is OpenAI-compatible.
The tradeoffs are model ceiling and throttling. Only open-weight models are offered, so quality tops out below frontier for hard reasoning, and the free tier's 6,000 tokens/min means one long-context request can eat a whole minute. Prompt caching gives 50% off cached input and a Batch API another 50%.
Groq vs its main rivals
The tools people actually weigh against Groq: Cerebras, Fireworks AI, Together AI. Same criteria for every column, including where Groq loses.
| CCerebras | FFireworks AI | Together AI | ||
|---|---|---|---|---|
Pricing & access | ||||
| Free tier | YesNo-card free tier | YesFree tier | YesSignup credits | No$5 min, no credits |
| Starting price | $0.05 /1M in | ~$0.10 /1M in | ~$0.10 /1M in | $0.17 /1M in |
Performance | ||||
| Peak throughput | ~1,000 tok/s | ~2,000 tok/s | Fast (GPU) | Fast (GPU) |
Models & capability | ||||
| Number of models | Curated (dozens) | Small set | 100s | 200+ |
| Frontier closed models | NoOpen only | NoOpen only | NoOpen only | NoOpen only |
| Fine-tuning | NoNo | UnknownLimited | YesYes | YesYes |
Scaling & limits | ||||
| Free-tier rate limits | 6K tok/min | Limited | Unknown | None |
| Prompt caching discount | Yes50% off | UnknownUnknown | YesYes | YesYes |
| Dedicated endpoints / GPUs | YesEnterprise | YesYes | YesYes | YesH100 $3.99/hr |
Developer & API | ||||
| REST API access | Free | Free | Free | Paid |
| OpenAI-compatible endpoint | YesYes | YesYes | YesYes | YesYes |
- Cheapest tokens in this comparison: Llama 3.1 8B at $0.05/$0.08 per 1M tokens
- High throughput (up to ~1,000 tokens/sec on GPT OSS 20B) for low-latency chat UIs
- No-card free tier is enough to ship a small production feature, not just a demo
- Open-weight models only, so quality tops out below frontier for hard reasoning
- Free tier throttles hard: 6,000 tokens/min means one long-context request can eat a whole minute
- Peak throughput trails Cerebras, and the model catalog is narrower than Together or Fireworks
Pay less for it
4 ways foundShip a small production feature within the free rate limits at $0
50% off cached input for repeated system prompts
50% off with a 24-hour to 7-day window for non-urgent jobs (does not stack with caching)
Use the $0.05/$0.08 tier for simple, high-volume tasks
StackTracker tracks what you actually pay for Groq and every other tool, flags overpayment, and shows the dollars you would save by switching.
Is it the right tool for you
- Your feature works on open models and response speed is part of the product
- You want the lowest cost per token in this comparison
- You want a no-card free tier that can carry a small production feature
- You need frontier closed models for hard reasoning → use OpenAI or Anthropic Claude
- You need a very wide open-model catalog plus fine-tuning → use Together AI or Fireworks AI
Track what Groq and the rest of your stack cost
StackTracker adds up every subscription, plus your hours, so you see the real number.
Prices and limits last verified 2026-07-20.
