Together AI logo

Together AI

An inference platform hosting 200+ open models behind one API, with fine-tuning and dedicated GPUs for scaling.

$0.17-$1.74 /1M input depending on model (Qwen 3.5 9B to DeepSeek V4 Pro), output $0.25-$3.48

What it actually does

Together AI serves 200+ open models (DeepSeek, Qwen, Kimi, Llama, and others) through one API, often the day they release. Serverless per-token pricing runs $0.17-$1.74 input depending on model, and the platform adds fine-tuning and dedicated GPU endpoints when you outgrow serverless.

The value is breadth and a scaling path, not the lowest per-token price. Fine-tuning starts at $0.48 per 1M training tokens to own a small custom model, and you can move from serverless tokens to dedicated endpoints and H100 GPUs at $3.99/hr as you grow.

There is no real free tier: the $25 signup credit was retired in 2025, leaving a fully prepaid platform with a $5 minimum top-up. Serverless rates often run above rivals for the same open models, for example Llama 3.3 70B at $1.04/$1.04 here versus $0.59/$0.79 on Groq.

Together AI vs its main rivals

The tools people actually weigh against Together AI: Fireworks AI, Groq, Replicate. Same criteria for every column, including where Together AI loses.

Together AI logoTogether AIFFireworks AIGroq logoGroqReplicate logoReplicate
Pricing & access
Free tierNo$5 min, no creditsYesSignup creditsYesNo-card free tierNoPay-as-you-go
Starting price$0.17 /1M in~$0.10 /1M in$0.05 /1M inPer-run / per-sec
Models & capability
Number of models200+100sCurated (dozens)1000s (community)
Fine-tuningYes$0.48/1M trainYesYesNoNoYesYes
Frontier closed modelsNoOpen onlyNoOpen onlyNoOpen onlyNoOpen only
New models day-oneYesYesYesYesNoSelectiveYesCommunity
Scaling & limits
Per-token price vs rivalsHigherLowerLowestVaries
Dedicated GPUs on demandH100 $3.99/hrYesEnterpriseYes
Prompt caching discountYesYesYesYesYes50% offUnknownUnknown
Developer & API
REST API accessPaidFreeFreePaid
OpenAI-compatible endpointYesYesYesYesYesYesNoOwn API
Worth it for
  • 200+ open models behind one API, including DeepSeek, Qwen, Kimi, and Llama the day they drop
  • Fine-tuning from $0.48 per 1M training tokens lets you own a small custom model cheaply
  • Clear upgrade path from serverless tokens to dedicated endpoints and H100 GPUs at $3.99/hr
Watch out for
  • Per-token rates run above rivals for the same open models (Llama 3.3 70B is $1.04/$1.04 vs $0.59/$0.79 on Groq)
  • No real free tier and a $5 minimum purchase, worse for testing than Groq or Gemini
  • Provisioned throughput bills per PTU-minute even when idle

Pay less for it

5 ways found
Route to smallest model

Serverless rates start at $0.17/$0.25 on the cheapest supported open models

Fine-tune to cut tokens

A small custom model from $0.48/1M training tokens can shrink per-request token counts

Batch inference

Batch endpoints discount non-urgent jobs

Dedicated endpoints at scale

H100 GPUs at $3.99/hr can beat per-token pricing at high, steady volume

Cheaper swap

GroqHosts many of the same open models at meaningfully lower per-token prices plus a no-card free tier, if you do not need Together's fine-tuning or 200+ catalog

See this for your whole stack, with your real numbers

StackTracker tracks what you actually pay for Together AI and every other tool, flags overpayment, and shows the dollars you would save by switching.

See plans

Is it the right tool for you

Pick it if
  • You need many open models plus fine-tuning under one roof
  • You expect to scale from serverless to dedicated capacity
  • You want new open models available on day one
Skip it if
  • You want the lowest per-token price on common open models → use Groq or Fireworks AI
  • You just want a free tier to test → use Groq or Google Gemini

Track what Together AI and the rest of your stack cost

StackTracker adds up every subscription, plus your hours, so you see the real number.

Prices and limits last verified 2026-07-20.