Pricing database 核验日期 2026-09-17
Every price below was read from the provider's official pricing page or models API on 2026-09-17. Unverifiable numbers are marked 待核验 rather than guessed. Prices change without notice — always confirm on the official page before committing your pipeline.
| Provider | Model | API model ID | Input $/1M | Cached $/1M | Output $/1M | Context | Notes | Source |
|---|---|---|---|---|---|---|---|---|
| OpenAI | GPT-6 Astra | gpt-6-astra | $10 | $1 | $50 | 1,050,000 | Flagship (knowledge cutoff 2026-04-30). Long-context surcharge: prompts >272K billed 2x input & cache / 1.5x output ($20 / $2 / $25 writes / $75). Batch & Flex −50%; fast mode 2x. Model docs: https://developers.openai.com/api/docs/models/gpt-6-astra.md | official ↗ |
| OpenAI | GPT-5.6-Luna | gpt-5.6-luna | $0.2 | $0.02 | $1.2 | 1,050,000 | Long-context surcharge: prompts >272K are billed 2x input / 1.5x output. Priority/flex tiers differ. | official ↗ |
| OpenAI | GPT-5 | gpt-5 | $1.25 | $0.125 | $10 | 400K | — | official ↗ |
| OpenAI | GPT-5 Mini | gpt-5-mini | $0.25 | $0.025 | $2 | 400K | — | official ↗ |
| OpenAI | GPT-5.2 | gpt-5.2 | $1.75 | $0.175 | $14 | 400K | — | official ↗ |
| Anthropic | Claude Sonnet 5 | claude-sonnet-5 | $2 | $0.2 | $10 | 1M | $2/$10 is now the standard price. Batch API −50%. | official ↗ |
| Anthropic | Claude Haiku 4.5 | claude-haiku-4.5 | $1 | $0.1 | $5 | 200K | Batch API −50%. | official ↗ |
| Anthropic | Claude Opus 5 | claude-opus-5 | $5 | $0.5 | $25 | 200K | Long-context (>200K prompt) billed $10/$50. | official ↗ |
| Anthropic | Claude Fable 5.1 | claude-fable-5.1 | $10 | $0.25 | $50 | 1M | Cache reads at 0.025× input price. | official ↗ |
| DeepSeek | DeepSeek V4.1 Flash | deepseek-flash | $0.3 | $0.006 | $1.2 | 1M | Peak rate. Off-peak $0.15/$0.60 (cache-hit $0.003). Anthropic-format endpoint available. | official ↗ |
| DeepSeek | DeepSeek V4 Pro | deepseek-v4-pro | $1.32 | $0.044 | $3.96 | 1M | Peak rate. Off-peak $0.66/$1.98. Service continued after 2026-09-14 per official notice. | official ↗ |
| Z.ai (Zhipu GLM) | GLM-5.3 | glm-5.3 | $1.4 | $0.26 | $4.4 | 1.31M | Flagship. Cached-input storage is limited-time free. | official ↗ |
| Z.ai (Zhipu GLM) | GLM-5.3-Flash | glm-5.3-flash | $0.15 | $0.03 | $0.5 | 1.31M | — | official ↗ |
| Z.ai (Zhipu GLM) | GLM-4.7 | glm-4.7 | $0.6 | $0.11 | $2.2 | 204K | — | official ↗ |
| Z.ai (Zhipu GLM) | GLM-4.7-Flash | glm-4.7-flash | Free | Free | Free | 204K | Free tier (officially listed at $0). | official ↗ |
| Z.ai (Zhipu GLM) | GLM-5 | glm-5 | $1 | $0.2 | $3.2 | 200K | — | official ↗ |
| Moonshot AI (Kimi) | Kimi K3 | kimi-k3 | $3 | $0.3 | $15 | 1,048,576 | — | official ↗ |
| Moonshot AI (Kimi) | Kimi K2.7 Code | kimi-k2.7-code | $0.95 | $0.19 | $4 | 262,144 | HighSpeed variant: $1.90/$8.00. | official ↗ |
| Moonshot AI (Kimi) | Kimi K2.6 | kimi-k2.6 | $0.95 | $0.16 | $4 | 262,144 | — | official ↗ |
| Alibaba Model Studio (Qwen) | Qwen3.8-Max | qwen3.8-max | $2 | 待核验 | $6 | 1M | qwen3.8-max-0902 same price. International (Model Studio) tier. | official ↗ |
| Alibaba Model Studio (Qwen) | Qwen3.8-Flash | qwen3.8-flash | $0.15 | 待核验 | $0.47 | 1M | — | official ↗ |
| Alibaba Model Studio (Qwen) | Qwen3.7-Plus | qwen3.7-plus | $0.4 | 待核验 | $1.6 | 1M | List price, 0–256K tier (20% limited-time discount may apply). >256K: $1.20/$4.80. | official ↗ |
| xAI (Grok) | Grok 4.6 | grok-4.6 | $2 | $0.5 | $6 | 500K | <200K prompt rate. ≥200K: $4/$12. | official ↗ |
| xAI (Grok) | Grok 4.3 | grok-4.3 | $1.25 | $0.2 | $2.5 | 1M | <200K prompt rate. ≥200K: $2.50/$5. | official ↗ |
| xAI (Grok) | Grok Build 0.1 | grok-build-0.1 | $1 | $0.2 | $2 | 256K | Coding-focused. <200K prompt rate; ≥200K: $2/$4. | official ↗ |
| SiliconFlow | DeepSeek-V4-Flash | DeepSeek-V4-Flash | $0.13 | $0.028 | $0.28 | 1,049K | Cheapest non-aggregator DeepSeek V4 Flash we verified (OpenRouter lists it at $0.07/$0.14). | official ↗ |
| SiliconFlow | GLM-5.3 | GLM-5.3 | $1.4 | $0.26 | $4.4 | 1,049K | Matches Z.ai first-party price. | official ↗ |
| SiliconFlow | GLM-5.3-Flash | GLM-5.3-Flash | $0.15 | $0.03 | $0.5 | 1,049K | — | official ↗ |
| SiliconFlow | Kimi-K2.6 | Kimi-K2.6 | $0.77 | $0.14 | $3.4 | 262K | Undercuts Moonshot first-party ($0.95/$4.00). | official ↗ |
| SiliconFlow | Kimi-K3 | Kimi-K3 | $2.7 | $0.27 | $13.5 | 1,049K | 10% under Moonshot first-party. | official ↗ |
| SiliconFlow | MiniMax-M2.5 | MiniMax-M2.5 | $0.3 | $0.03 | $1.2 | 197K | — | official ↗ |
| OpenRouter | DeepSeek V4 Flash (deepseek-v4-flash) | deepseek/deepseek-v4-flash | $0.07 | $0.014 | $0.14 | 1,048,576 | Cheapest DeepSeek-class listing we verified. Aggregator price, can drift per provider. | official ↗ |
| OpenRouter | DeepSeek V3.2 (deepseek-chat) | deepseek/deepseek-chat | $0.2574 | 待核验 | $1.0287 | 163,840 | Aggregator price from OpenRouter's public API. | official ↗ |
| OpenRouter | GLM-4.6 | z-ai/glm-4.6 | $0.43 | $0.08 | $1.75 | 204,800 | — | official ↗ |
| OpenRouter | Llama 4 Maverick (Meta) | meta-llama/llama-4-maverick | $0.1875 | 待核验 | $0.6525 | 1,048,576 | — | official ↗ |
| OpenRouter | Llama 3.3 70B (Meta) | meta-llama/llama-3.3-70b-instruct | $0.1 | 待核验 | $0.32 | 131,072 | — | official ↗ |
| OpenRouter | GPT-5 | openai/gpt-5 | $1.25 | $0.125 | $10 | 400,000 | Same as OpenAI first-party. | official ↗ |
| OpenRouter | Claude Sonnet 5 | anthropic/claude-sonnet-5 | $2 | $0.2 | $10 | 1,000,000 | Same as Anthropic first-party. | official ↗ |
| OpenRouter | Qwen3.7-Plus | qwen/qwen3.7-plus | $0.32 | $0.064 | $1.28 | 1,000,000 | Under Alibaba first-party list price ($0.40/$1.60). | official ↗ |
Source pages
- OpenAI — https://developers.openai.com/api/docs/pricing
- Anthropic — https://platform.claude.com/docs/en/about-claude/pricing
- DeepSeek — https://api-docs.deepseek.com/quick_start/pricing/
- Z.ai (Zhipu GLM) — https://docs.z.ai/guides/overview/pricing
- Moonshot AI (Kimi) — https://platform.kimi.ai/docs/pricing/chat
- Alibaba Cloud Model Studio (Qwen) — https://www.alibabacloud.com/help/en/model-studio/model-pricing
- xAI (Grok) — https://docs.x.ai/docs/models
- SiliconFlow — https://siliconflow.com/pricing
- OpenRouter — https://openrouter.ai/models
Methodology & caveats
- DeepSeek publishes off-peak rates (01:00–04:00 and 06:00–10:00 UTC, Monday–Friday) at ½ peak. We list peak; divide by 2 for off-peak. Cache-hit input for V4.1 Flash is $0.006 peak / $0.003 off-peak.
- xAI, OpenAI, Anthropic apply long-context surcharges above a prompt threshold; the table shows the standard tier (see Notes column).
- OpenRouter prices are aggregator pass-through and can drift per-provider; we snapshot the default listed price from their public models API.
- OpenAI cached-input prices apply to automatic prompt caching; batch API typically −50% on both directions.
- Qwen cache-read rates on Alibaba's page are tiered by token count and model version — marked 待核验 here rather than approximated.
Next step: put your own numbers in the monthly cost calculator →