OpenAI
o4-mini API Pricing
Page notes it has been "succeeded by GPT-5 mini"; snapshot o4-mini-2025-04-16 marked deprecated. Also has separate fine-tuning-tier pricing: standard inference $4.00/$16.00 (cached $1.00) per MTok, batch $2.00/$8.00, training $100/hour.
Verified against OpenAI’s pricing page on 2026-07-02.
View source ↗Facts
Specs
- Familyo-series
- API IDo4-mini
- ModalityText
- Context window200K tokens
- Max output—
- Release dateNot published
Pricing (per 1M tokens)
- Input$1.10 / 1M tokens
- Output$4.40 / 1M tokens
- Cached input$0.275 / 1M tokens
- Cache write—
- Batch discount—
Capabilities
reasoning
Notes
Page notes it has been "succeeded by GPT-5 mini"; snapshot o4-mini-2025-04-16 marked deprecated. Also has separate fine-tuning-tier pricing: standard inference $4.00/$16.00 (cached $1.00) per MTok, batch $2.00/$8.00, training $100/hour.
Worked cost examples
Estimated monthly cost on three common usage shapes
| Workload | Est. monthly cost |
|---|---|
Chatbot 800 input / 200 output tokens per request, 100,000 requests/month | $176.00 |
RAG pipeline 6,000 input / 700 output tokens per request, 30,000 requests/month | $290.40 |
Batch summarization 20,000 input / 1,500 output tokens per request, 5,000 requests/month | $143.00 |
Price history
Changes we have documented since we began tracking
No price changes recorded since we began tracking.
o4-mini pricing FAQ
- How much does o4-mini cost per 1M tokens?
- o4-mini costs $1.10 per 1M input tokens and $4.40 per 1M output tokens.
- How much does o4-mini cost per request?
- A typical request of 800 input and 200 output tokens costs about $0.0018 on o4-mini.
- Is o4-mini cheaper than o1?
- o4-mini is cheaper than o1 on input tokens, by about 93%.
- Does o4-mini have cache pricing?
- Yes — cached input reads cost $0.275 per 1M tokens.