LLM Cost Per Million Tokens: A Full Provider Comparison

Compare LLM cost per million tokens across OpenAI, Anthropic, Google, and DeepSeek in 2026 — input, output, cached input, and batch pricing in one table.

LayerFlow Team8 min read
LLM Cost Per Million Tokens: A Full Provider Comparison — LayerFlow blog illustration

Comparing LLM providers by sticker price is easy once you standardize on one unit: US dollars per 1 million tokens. This guide walks through representative 2026 rates for input, output, cached input, and batch, with the caveat that exact numbers change quarterly.

The patterns matter more than the exact prices, and those patterns have been stable for two years.

Typical 2026 rates per million tokens

  • OpenAI GPT-4.1 / o-series: ~$2-4 input, ~$8-16 output; cached input ~$0.25-1.
  • Anthropic Claude: ~$3 input, ~$15 output; cached input ~90% cheaper.
  • Google Gemini Pro: ~$1.25-2.50 input, ~$10 output; Flash cheaper.
  • DeepSeek: among the cheapest — well under $1 input, few dollars output.
  • Batch (async) endpoints: typically 50% off on both directions.

How to read these numbers

Output tokens cost 3-8x input, so a 'cheap' model with verbose output can still be expensive. Caching stable prefixes (system prompts, docs) slashes effective input cost. Always model your actual mix rather than comparing a single price.

Automate the comparison

The sane way to compare for your workload is to route a sample of real traffic across providers in a gateway and read the cost-per-1M breakdown from analytics. That turns pricing debates into measured data.

FAQ

How much does 1 million tokens cost across providers?+

Roughly $1-4 input and $8-16 output for frontier models in 2026, with cheaper tiers like Gemini Flash and DeepSeek under $1 input.

Which LLM is cheapest per token?+

DeepSeek and small-tier models (mini/flash/small) are typically the cheapest per token, followed by Google's Flash options.

Does cached input save much?+

Yes — cached input is often 50-90% cheaper than fresh input, so caching stable prefixes is one of the biggest savings levers.

Related posts

LayerFlow

Try the AI workspace

Save prompts, compare models, and set hard budgets in one place.