AI API Pricing Tracker 2026
Compare LLM API pricing across 5 providers and 14 models. All prices shown per 1 million tokens. Prices change frequently — verify with the provider before budgeting.
OpenAI
Last verified: August 2026| Model | Tier | Input / 1M | Output / 1M | Context | Notes |
|---|---|---|---|---|---|
| GPT-4obest-overall | Frontier | $2.50 | $10.00 | 128K | Vision + audio. Most capable OpenAI multimodal model. |
| o3-mini | Reasoning | $1.10 | $4.40 | 200K | Reasoning model. Best for coding, math, structured analysis. |
| GPT-4o minibest-value | Efficient | $0.15 | $0.60 | 128K | Best price-performance for most product use cases. |
| GPT-3.5 Turbo | Budget | $0.50 | $1.50 | 16K | Legacy. Use GPT-4o mini instead for better results at lower cost. |
Anthropic
Last verified: August 2026| Model | Tier | Input / 1M | Output / 1M | Context | Notes |
|---|---|---|---|---|---|
| Claude 3.5 Sonnetbest-coding | Frontier | $3.00 | $15.00 | 200K | Best for coding and long-context analysis. Top coding benchmark. |
| Claude 3 Opus | Frontier | $15.00 | $75.00 | 200K | Anthropic's most powerful. Very high cost — use selectively. |
| Claude 3.5 Haiku | Efficient | $0.80 | $4.00 | 200K | Fast, affordable. Great for classification and extraction. |
| Model | Tier | Input / 1M | Output / 1M | Context | Notes |
|---|---|---|---|---|---|
| Gemini 1.5 Prolongest-context | Frontier | $3.50 | $10.50 | 2M | Largest context window (2M tokens). Best for very long documents. |
| Gemini 1.5 Flashcheapest | Efficient | $0.075 | $0.30 | 1M | Extremely cheap. Best cost-per-token for high-volume applications. |
| Gemini 2.0 Flash | Efficient | $0.10 | $0.40 | 1M | Next-gen Flash. Better reasoning than 1.5 Flash at similar cost. |
Meta (via API)
Last verified: August 2026| Model | Tier | Input / 1M | Output / 1M | Context | Notes |
|---|---|---|---|---|---|
| Llama 3.1 70B | Open Source | $0.35 | $0.40 | 128K | Most capable open-weights model. Price varies by provider. |
| Llama 3.1 8B | Open Source | $0.06 | $0.06 | 128K | Ultra-cheap for simple tasks. Can self-host at near-zero cost. |
Mistral
Last verified: August 2026| Model | Tier | Input / 1M | Output / 1M | Context | Notes |
|---|---|---|---|---|---|
| Mistral Large 2 | Frontier | $2.00 | $6.00 | 128K | Strong European alternative. Good multilingual support. |
| Mistral 7B | Budget | $0.25 | $0.25 | 32K | Very fast and cheap. Best for simple classification/extraction. |
Quick Decision Guide
Best overall qualityGPT-4o or Claude 3.5 Sonnet
Best price/performanceGPT-4o mini or Gemini 1.5 Flash
Coding and technical tasksClaude 3.5 Sonnet
Very long documents (>100K tokens)Gemini 1.5 Pro (2M context)
Lowest possible costGemini 1.5 Flash or Llama 3.1 8B (self-hosted)
Reasoning / math / scienceo3-mini
Open source / self-hostableLlama 3.1 70B
Disclaimer: Prices are approximate and change frequently. Always verify current pricing on the provider's official pricing page before production deployment. Prices may vary by volume, region, or enterprise contract.