Directory
AI model API pricing
Every current, legacy and deprecated text and multimodal model we track across 11 providers, with input, output and cached token rates in USD per 1M tokens. Prices checked 2026-07-02.
- Amazon Nova 2.0 Lite
Amazon (Nova)
- Context—
- Input$0.33 / 1M tokens
- Output$2.75 / 1M tokens
- Cached$0.0825 / 1M tokens
- Amazon Nova 2.0 Omni
Amazon (Nova)
- Context—
- Input$0.30 / 1M tokens
- Output$2.80 / 1M tokens
- Cached—
- Amazon Nova 2.0 Pro
Amazon (Nova)
- Context—
- Input$1.38 / 1M tokens
- Output$11.00 / 1M tokens
- Cached—
- Amazon Nova Lite
Amazon (Nova)
- Context—
- Input$0.06 / 1M tokens
- Output$0.24 / 1M tokens
- Cached$0.015 / 1M tokens
- Amazon Nova Micro
Amazon (Nova)
- Context—
- Input$0.035 / 1M tokens
- Output$0.14 / 1M tokens
- Cached$0.0088 / 1M tokens
- Amazon Nova Premier
Amazon (Nova)
- Context—
- Input$2.50 / 1M tokens
- Output$12.50 / 1M tokens
- Cached$0.625 / 1M tokens
- Amazon Nova Pro
Amazon (Nova)
- Context—
- Input$0.80 / 1M tokens
- Output$3.20 / 1M tokens
- Cached$0.20 / 1M tokens
- Amazon Nova Pro (Latency Optimized)
Amazon (Nova)
- Context—
- Input$1.00 / 1M tokens
- Output$4.00 / 1M tokens
- Cached—
- Aya Expanse 32B
Cohere
- Context—
- Input$0.50 / 1M tokens
- Output$1.50 / 1M tokens
- Cached—
- Aya Expanse 8B
Cohere
- Context—
- Input$0.50 / 1M tokens
- Output$1.50 / 1M tokens
- Cached—
- chat-latest
OpenAI
- Context400K
- Input$5.00 / 1M tokens
- Output$30.00 / 1M tokens
- Cached$0.50 / 1M tokens
- Claude Fable 5
Anthropic
- Context1M
- Input$10.00 / 1M tokens
- Output$50.00 / 1M tokens
- Cached$1.00 / 1M tokens
- Claude Haiku 4.5
Anthropic
- Context200K
- Input$1.00 / 1M tokens
- Output$5.00 / 1M tokens
- Cached$0.10 / 1M tokens
- Claude Mythos 5
Anthropic
- Context1M
- Input$10.00 / 1M tokens
- Output$50.00 / 1M tokens
- Cached$1.00 / 1M tokens
- Claude Opus 4.8
Anthropic
- Context1M
- Input$5.00 / 1M tokens
- Output$25.00 / 1M tokens
- Cached$0.50 / 1M tokens
- Claude Sonnet 5
Anthropic
- Context1M
- Input$2.00 / 1M tokens
- Output$10.00 / 1M tokens
- Cached$0.20 / 1M tokens
- Codestral
Mistral AI
- Context—
- Input$0.30 / 1M tokens
- Output$0.90 / 1M tokens
- Cached—
- Command AUnverified
Cohere
- Context256K
- Input—
- Output—
- Cached—
- Command A ReasoningUnverified
Cohere
- Context256K
- Input—
- Output—
- Cached—
- Command A TranslateUnverified
Cohere
- Context8K
- Input—
- Output—
- Cached—
- Command A VisionUnverified
Cohere
- Context128K
- Input—
- Output—
- Cached—
- Command A+Unverified
Cohere
- Context128K
- Input—
- Output—
- Cached—
- Command R (08-2024)Unverified
Cohere
- Context128K
- Input—
- Output—
- Cached—
- Command R7BUnverified
Cohere
- Context128K
- Input—
- Output—
- Cached—
- DeepSeek V4 Flash
DeepSeek
- Context1M
- Input$0.14 / 1M tokens
- Output$0.28 / 1M tokens
- Cached$0.0028 / 1M tokens
- DeepSeek V4 Pro
DeepSeek
- Context1M
- Input$0.435 / 1M tokens
- Output$0.87 / 1M tokens
- Cached$0.0036 / 1M tokens
- Devstral 2
Mistral AI
- Context—
- Input$0.40 / 1M tokens
- Output$2.00 / 1M tokens
- Cached—
- Devstral Small 2
Mistral AI
- Context—
- Input$0.10 / 1M tokens
- Output$0.30 / 1M tokens
- Cached—
- Context—
- Input$1.25 / 1M tokens
- Output$10.00 / 1M tokens
- Cached—
- Gemini 2.5 Flash
Google
- Context—
- Input$0.30 / 1M tokens
- Output$2.50 / 1M tokens
- Cached$0.03 / 1M tokens
- Gemini 2.5 Flash-Lite
Google
- Context—
- Input$0.10 / 1M tokens
- Output$0.40 / 1M tokens
- Cached$0.01 / 1M tokens
- Gemini 2.5 Pro
Google
- Context—
- Input$1.25 / 1M tokens
- Output$10.00 / 1M tokens
- Cached$0.125 / 1M tokens
- Gemini 3 Flash Preview
Google
- Context—
- Input$0.50 / 1M tokens
- Output$3.00 / 1M tokens
- Cached$0.05 / 1M tokens
- Gemini 3.1 Flash-Lite
Google
- Context—
- Input$0.25 / 1M tokens
- Output$1.50 / 1M tokens
- Cached$0.025 / 1M tokens
- Gemini 3.1 Pro Preview
Google
- Context—
- Input$2.00 / 1M tokens
- Output$12.00 / 1M tokens
- Cached$0.20 / 1M tokens
- Gemini 3.5 Flash
Google
- Context1M
- Input$1.50 / 1M tokens
- Output$9.00 / 1M tokens
- Cached$0.15 / 1M tokens
- GPT-5.3-Codex
OpenAI
- Context400K
- Input$1.75 / 1M tokens
- Output$14.00 / 1M tokens
- Cached$0.175 / 1M tokens
- GPT-5.4
OpenAI
- Context1.1M
- Input$2.50 / 1M tokens
- Output$15.00 / 1M tokens
- Cached$0.25 / 1M tokens
- GPT-5.4 mini
OpenAI
- Context400K
- Input$0.75 / 1M tokens
- Output$4.50 / 1M tokens
- Cached$0.075 / 1M tokens
- GPT-5.4 nano
OpenAI
- Context400K
- Input$0.20 / 1M tokens
- Output$1.25 / 1M tokens
- Cached$0.02 / 1M tokens
- GPT-5.5
OpenAI
- Context1.1M
- Input$5.00 / 1M tokens
- Output$30.00 / 1M tokens
- Cached$0.50 / 1M tokens
- GPT-5.5 Pro
OpenAI
- Context1.1M
- Input$30.00 / 1M tokens
- Output$180.00 / 1M tokens
- Cached—
- Context1M
- Input$1.25 / 1M tokens
- Output$2.50 / 1M tokens
- Cached$0.20 / 1M tokens
- Context1M
- Input$1.25 / 1M tokens
- Output$2.50 / 1M tokens
- Cached$0.20 / 1M tokens
- Context1M
- Input$1.25 / 1M tokens
- Output$2.50 / 1M tokens
- Cached$0.20 / 1M tokens
- Grok 4.3
xAI
- Context1M
- Input$1.25 / 1M tokens
- Output$2.50 / 1M tokens
- Cached$0.20 / 1M tokens
- Context256K
- Input$1.00 / 1M tokens
- Output$2.00 / 1M tokens
- Cached—
- Leanstral
Mistral AI
- Context—
- Input$0.00 / 1M tokens
- Output$0.00 / 1M tokens
- Cached—
- Llama 4 Maverick
Meta
- Context1M
- Input$0.15 / 1M tokens
- Output$0.60 / 1M tokens
- Cached—
- Llama 4 Scout
Meta
- Context10M
- Input$0.10 / 1M tokens
- Output$0.30 / 1M tokens
- Cached—
- Magistral Medium
Mistral AI
- Context—
- Input$2.00 / 1M tokens
- Output$5.00 / 1M tokens
- Cached—
- Magistral Small
Mistral AI
- Context—
- Input$0.50 / 1M tokens
- Output$1.50 / 1M tokens
- Cached—
- Ministral 3 - 14B
Mistral AI
- Context—
- Input$0.20 / 1M tokens
- Output$0.20 / 1M tokens
- Cached—
- Ministral 3 - 3B
Mistral AI
- Context—
- Input$0.10 / 1M tokens
- Output$0.10 / 1M tokens
- Cached—
- Ministral 3 - 8B
Mistral AI
- Context—
- Input$0.15 / 1M tokens
- Output$0.15 / 1M tokens
- Cached—
- Mistral Large 3
Mistral AI
- Context—
- Input$0.50 / 1M tokens
- Output$1.50 / 1M tokens
- Cached—
- Mistral Medium 3.5
Mistral AI
- Context—
- Input$1.50 / 1M tokens
- Output$7.50 / 1M tokens
- Cached—
- Mistral Moderation
Mistral AI
- Context—
- Input$0.10 / 1M tokens
- Output—
- Cached—
- Mistral OCR 4
Mistral AI
- Context—
- Input—
- Output—
- Cached—
- Mistral Small 4
Mistral AI
- Context—
- Input$0.15 / 1M tokens
- Output$0.60 / 1M tokens
- Cached—
- Qwen-Flash
Alibaba Cloud (Qwen)
- Context—
- Input$0.05 / 1M tokens
- Output$0.40 / 1M tokens
- Cached—
- Qwen-Long
Alibaba Cloud (Qwen)
- Context—
- Input$0.072 / 1M tokens
- Output$0.287 / 1M tokens
- Cached—
- Qwen-MT-Flash
Alibaba Cloud (Qwen)
- Context—
- Input$0.16 / 1M tokens
- Output$0.49 / 1M tokens
- Cached—
- Qwen-Plus
Alibaba Cloud (Qwen)
- Context—
- Input$0.115 / 1M tokens
- Output$1.15 / 1M tokens
- Cached—
- Qwen3-Max
Alibaba Cloud (Qwen)
- Context—
- Input$1.20 / 1M tokens
- Output$6.00 / 1M tokens
- Cached—
- Qwen3-Omni-FlashUnverified
Alibaba Cloud (Qwen)
- Context—
- Input—
- Output—
- Cached—
- Qwen3-VL-Plus
Alibaba Cloud (Qwen)
- Context—
- Input$0.20 / 1M tokens
- Output$1.60 / 1M tokens
- Cached—
- Qwen3.5-Flash
Alibaba Cloud (Qwen)
- Context—
- Input$0.10 / 1M tokens
- Output$0.40 / 1M tokens
- Cached—
- Qwen3.6-Flash
Alibaba Cloud (Qwen)
- Context—
- Input$0.25 / 1M tokens
- Output$1.50 / 1M tokens
- Cached—
- Qwen3.7-Max
Alibaba Cloud (Qwen)
- Context—
- Input$2.50 / 1M tokens
- Output$7.50 / 1M tokens
- Cached—
- Qwen3.7-Plus
Alibaba Cloud (Qwen)
- Context—
- Input$0.40 / 1M tokens
- Output$1.60 / 1M tokens
- Cached—
- QwQ-Plus
Alibaba Cloud (Qwen)
- Context—
- Input$0.80 / 1M tokens
- Output$2.40 / 1M tokens
- Cached—
- Rerank 3.5Unverified
Cohere
- Context4K
- Input—
- Output—
- Cached—
- Rerank English v3.0Unverified
Cohere
- Context4K
- Input—
- Output—
- Cached—
- Rerank Multilingual v3.0Unverified
Cohere
- Context4K
- Input—
- Output—
- Cached—
- Rerank v4.0 FastUnverified
Cohere
- Context32K
- Input—
- Output—
- Cached—
- Rerank v4.0 ProUnverified
Cohere
- Context32K
- Input—
- Output—
- Cached—
- rerank-2.5
Voyage AI
- Context—
- Input$0.05 / 1M tokens
- Output—
- Cached—
- rerank-2.5-lite
Voyage AI
- Context—
- Input$0.02 / 1M tokens
- Output—
- Cached—
- Claude Opus 4.5Legacy
Anthropic
- Context200K
- Input$5.00 / 1M tokens
- Output$25.00 / 1M tokens
- Cached$0.50 / 1M tokens
- Claude Opus 4.6Legacy
Anthropic
- Context1M
- Input$5.00 / 1M tokens
- Output$25.00 / 1M tokens
- Cached$0.50 / 1M tokens
- Claude Opus 4.7Legacy
Anthropic
- Context1M
- Input$5.00 / 1M tokens
- Output$25.00 / 1M tokens
- Cached$0.50 / 1M tokens
- Claude Sonnet 4.5Legacy
Anthropic
- Context200K
- Input$3.00 / 1M tokens
- Output$15.00 / 1M tokens
- Cached$0.30 / 1M tokens
- Claude Sonnet 4.6Legacy
Anthropic
- Context1M
- Input$3.00 / 1M tokens
- Output$15.00 / 1M tokens
- Cached$0.30 / 1M tokens
- Command R+ (08-2024)Legacy
Cohere
- Context128K
- Input$2.50 / 1M tokens
- Output$10.00 / 1M tokens
- Cached—
- GPT-4.1Legacy
OpenAI
- Context1M
- Input$2.00 / 1M tokens
- Output$8.00 / 1M tokens
- Cached$0.50 / 1M tokens
- GPT-4.1 miniLegacy
OpenAI
- Context1M
- Input$0.40 / 1M tokens
- Output$1.60 / 1M tokens
- Cached$0.10 / 1M tokens
- GPT-4.1 nanoLegacy
OpenAI
- Context1M
- Input$0.10 / 1M tokens
- Output$0.40 / 1M tokens
- Cached$0.025 / 1M tokens
- GPT-4oLegacy
OpenAI
- Context128K
- Input$2.50 / 1M tokens
- Output$10.00 / 1M tokens
- Cached$1.25 / 1M tokens
- GPT-4o miniLegacy
OpenAI
- Context128K
- Input$0.15 / 1M tokens
- Output$0.60 / 1M tokens
- Cached$0.075 / 1M tokens
- GPT-5.4 ProLegacy
OpenAI
- Context1.1M
- Input$30.00 / 1M tokens
- Output$180.00 / 1M tokens
- Cached—
- Legacy
- Context131K
- Input$0.10 / 1M tokens
- Output$0.32 / 1M tokens
- Cached—
- Mistral NeMoLegacy
Mistral AI
- Context—
- Input$0.15 / 1M tokens
- Output$0.15 / 1M tokens
- Cached—
- Mixtral 8x22BLegacy
Mistral AI
- Context—
- Input$2.00 / 1M tokens
- Output$6.00 / 1M tokens
- Cached—
- Mixtral 8x7BLegacy
Mistral AI
- Context—
- Input$0.70 / 1M tokens
- Output$0.70 / 1M tokens
- Cached—
- o3Legacy
OpenAI
- Context200K
- Input$2.00 / 1M tokens
- Output$8.00 / 1M tokens
- Cached$0.50 / 1M tokens
- o3-deep-researchLegacy
OpenAI
- Context—
- Input—
- Output—
- Cached—
- o3-miniLegacy
OpenAI
- Context200K
- Input$1.10 / 1M tokens
- Output$4.40 / 1M tokens
- Cached$0.55 / 1M tokens
- o4-miniLegacy
OpenAI
- Context200K
- Input$1.10 / 1M tokens
- Output$4.40 / 1M tokens
- Cached$0.275 / 1M tokens
- o4-mini-deep-researchLegacy
OpenAI
- Context—
- Input—
- Output—
- Cached—
- Qwen-TurboLegacy
Alibaba Cloud (Qwen)
- Context—
- Input$0.05 / 1M tokens
- Output$0.20 / 1M tokens
- Cached—
- rerank-2Legacy
Voyage AI
- Context—
- Input$0.05 / 1M tokens
- Output—
- Cached—
- rerank-2-liteLegacy
Voyage AI
- Context—
- Input$0.02 / 1M tokens
- Output—
- Cached—
- Claude Haiku 3.5DeprecatedUnverified
Anthropic
- Context—
- Input$0.80 / 1M tokens
- Output$4.00 / 1M tokens
- Cached$0.08 / 1M tokens
- Claude Opus 4DeprecatedUnverified
Anthropic
- Context—
- Input$15.00 / 1M tokens
- Output$75.00 / 1M tokens
- Cached$1.50 / 1M tokens
- Claude Opus 4.1Deprecated
Anthropic
- Context200K
- Input$15.00 / 1M tokens
- Output$75.00 / 1M tokens
- Cached$1.50 / 1M tokens
- Claude Sonnet 4DeprecatedUnverified
Anthropic
- Context—
- Input$3.00 / 1M tokens
- Output$15.00 / 1M tokens
- Cached$0.30 / 1M tokens
- CommandDeprecated
Cohere
- Context—
- Input$1.00 / 1M tokens
- Output$2.00 / 1M tokens
- Cached—
- Command LightDeprecated
Cohere
- Context—
- Input$0.30 / 1M tokens
- Output$0.60 / 1M tokens
- Cached—
- Command R (03-2024)Deprecated
Cohere
- Context—
- Input$0.50 / 1M tokens
- Output$1.50 / 1M tokens
- Cached—
- Command R+ (04-2024)Deprecated
Cohere
- Context—
- Input$3.00 / 1M tokens
- Output$15.00 / 1M tokens
- Cached—
- Gemini 2.0 FlashDeprecated
Google
- Context—
- Input$0.10 / 1M tokens
- Output$0.40 / 1M tokens
- Cached—
- Gemini 2.0 Flash-LiteDeprecated
Google
- Context—
- Input$0.075 / 1M tokens
- Output$0.30 / 1M tokens
- Cached—
- o1Deprecated
OpenAI
- Context200K
- Input$15.00 / 1M tokens
- Output$60.00 / 1M tokens
- Cached$7.50 / 1M tokens
- rerank-1Deprecated
Voyage AI
- Context—
- Input$0.05 / 1M tokens
- Output—
- Cached—
- rerank-lite-1Deprecated
Voyage AI
- Context—
- Input$0.02 / 1M tokens
- Output—
- Cached—
Embedding models
Priced per input token only — there is no output leg for embeddings.
- Amazon Nova Multimodal Embeddings
Amazon (Nova)
- Context—
- Input$0.135 / 1M tokens
- Output—
- Cached—
- Codestral Embed
Mistral AI
- Context—
- Input$0.15 / 1M tokens
- Output—
- Cached—
- Embed English Light v3.0Unverified
Cohere
- Context512
- Input—
- Output—
- Cached—
- Embed English v3.0Unverified
Cohere
- Context512
- Input—
- Output—
- Cached—
- Unverified
- Context512
- Input—
- Output—
- Cached—
- Embed Multilingual v3.0Unverified
Cohere
- Context512
- Input—
- Output—
- Cached—
- Embed v4.0Unverified
Cohere
- Context128K
- Input—
- Output—
- Cached—
- Gemini Embedding 001Legacy
Google
- Context—
- Input$0.15 / 1M tokens
- Output—
- Cached—
- Gemini Embedding 2
Google
- Context—
- Input$0.20 / 1M tokens
- Output—
- Cached—
- Mistral Embed
Mistral AI
- Context—
- Input$0.10 / 1M tokens
- Output—
- Cached—
- Text Embedding v4Unverified
Alibaba Cloud (Qwen)
- Context—
- Input—
- Output—
- Cached—
- text-embedding-3-large
OpenAI
- Context—
- Input$0.13 / 1M tokens
- Output—
- Cached—
- text-embedding-3-small
OpenAI
- Context—
- Input$0.02 / 1M tokens
- Output—
- Cached—
- text-embedding-ada-002Legacy
OpenAI
- Context—
- Input$0.10 / 1M tokens
- Output—
- Cached—
- voyage-3Legacy
Voyage AI
- Context—
- Input$0.06 / 1M tokens
- Output—
- Cached—
- voyage-3-largeLegacy
Voyage AI
- Context—
- Input$0.18 / 1M tokens
- Output—
- Cached—
- voyage-3-liteLegacy
Voyage AI
- Context—
- Input$0.02 / 1M tokens
- Output—
- Cached—
- voyage-3.5Legacy
Voyage AI
- Context—
- Input$0.06 / 1M tokens
- Output—
- Cached—
- voyage-3.5-liteLegacy
Voyage AI
- Context—
- Input$0.02 / 1M tokens
- Output—
- Cached—
- voyage-4
Voyage AI
- Context—
- Input$0.06 / 1M tokens
- Output—
- Cached—
- voyage-4-large
Voyage AI
- Context—
- Input$0.12 / 1M tokens
- Output—
- Cached—
- voyage-4-lite
Voyage AI
- Context—
- Input$0.02 / 1M tokens
- Output—
- Cached—
- voyage-code-2Legacy
Voyage AI
- Context—
- Input$0.12 / 1M tokens
- Output—
- Cached—
- voyage-code-3
Voyage AI
- Context—
- Input$0.18 / 1M tokens
- Output—
- Cached—
- voyage-context-3
Voyage AI
- Context—
- Input$0.18 / 1M tokens
- Output—
- Cached—
- voyage-finance-2
Voyage AI
- Context—
- Input$0.12 / 1M tokens
- Output—
- Cached—
- voyage-large-2Legacy
Voyage AI
- Context—
- Input$0.12 / 1M tokens
- Output—
- Cached—
- voyage-law-2
Voyage AI
- Context—
- Input$0.12 / 1M tokens
- Output—
- Cached—
- voyage-multilingual-2
Voyage AI
- Context—
- Input$0.12 / 1M tokens
- Output—
- Cached—
- voyage-multimodal-3Legacy
Voyage AI
- Context—
- Input$0.12 / 1M tokens
- Output—
- Cached—
- voyage-multimodal-3.5
Voyage AI
- Context—
- Input$0.12 / 1M tokens
- Output—
- Cached—
Image & audio models (15)
Usually billed per image, per second or with separate speech/text rates rather than per token — see each model’s page for its documented unit price.
| Model | Provider | Context | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|---|---|
| Amazon Nova Canvas | Amazon (Nova) | — | — | — | — |
| Amazon Nova Reel | Amazon (Nova) | — | — | — | — |
| Amazon Nova Sonic | Amazon (Nova) | — | — | — | — |
| Amazon Nova Sonic 2.0 | Amazon (Nova) | — | — | — | — |
| Grok Imagine (Image, Quality) | xAI | — | — | — | — |
| Grok Imagine (Image, Standard) | xAI | — | — | — | — |
| Grok Imagine (Video 1.5) | xAI | — | — | — | — |
| Grok Imagine (Video) | xAI | — | — | — | — |
| Grok Speech to Text | xAI | — | — | — | — |
| Grok Text to Speech | xAI | — | — | — | — |
| Grok Voice (Realtime) | xAI | — | — | — | — |
| Voxtral Mini Transcribe 2 | Mistral AI | — | — | — | — |
| Voxtral Mini Transcribe Realtime | Mistral AI | — | — | — | — |
| Voxtral Small | Mistral AI | — | $0.10 | $0.40 | — |
| Voxtral TTS | Mistral AI | — | — | — | — |
- Amazon Nova Canvas
Amazon (Nova)
- Context—
- Input—
- Output—
- Cached—
- Amazon Nova Reel
Amazon (Nova)
- Context—
- Input—
- Output—
- Cached—
- Amazon Nova Sonic
Amazon (Nova)
- Context—
- Input—
- Output—
- Cached—
- Amazon Nova Sonic 2.0
Amazon (Nova)
- Context—
- Input—
- Output—
- Cached—
- Context—
- Input—
- Output—
- Cached—
- Context—
- Input—
- Output—
- Cached—
- Context—
- Input—
- Output—
- Cached—
- Context—
- Input—
- Output—
- Cached—
- Context—
- Input—
- Output—
- Cached—
- Context—
- Input—
- Output—
- Cached—
- Context—
- Input—
- Output—
- Cached—
- Voxtral Mini Transcribe 2
Mistral AI
- Context—
- Input—
- Output—
- Cached—
- Voxtral Mini Transcribe Realtime
Mistral AI
- Context—
- Input—
- Output—
- Cached—
- Voxtral Small
Mistral AI
- Context—
- Input$0.10 / 1M tokens
- Output$0.40 / 1M tokens
- Cached—
- Voxtral TTS
Mistral AI
- Context—
- Input—
- Output—
- Cached—