Models and pricing
Every model we serve, with the id you pass in your request. Rates are per 1M tokens, in US dollars.
Input and output are on screen; swipe the table sideways for the cache rates.
| Model | Context | Input | Output | Cache read | Cache write |
|---|---|---|---|---|---|
| GLM | |||||
GLM-5.3glm-5.31M context | 1M | $1.40 | $4.40 | $0.26 | — |
GLM-5.2glm-5.21M context | 1M | $1.05 | $3.30 | $0.195 | — |
GLM-5.1glm-5.1202K context | 202K | $1.05 | $3.30 | $0.195 | — |
| Kimi | |||||
Kimi K3kimi-k31M context | 1M | $3.00 | $15.00 | $0.30 | — |
Kimi K2.7 Codekimi-k2.7-code262K context | 262K | $0.7125 | $3.00 | $0.1425 | — |
Kimi K2.6kimi-k2.6262K context | 262K | $0.7125 | $3.00 | $0.12 | — |
| MiMo | |||||
MiMo V2.5 Promimo-v2.5-pro1M context | 1M | $0.435 | $0.87 | $0.003625 | — |
MiMo V2.5mimo-v2.51M context | 1M | $0.105 | $0.21 | $0.0021 | — |
| MiniMax | |||||
MiniMax M3minimax-m31M context | 1M | $0.225 | $0.90 | $0.045 | — |
MiniMax M2.7minimax-m2.7204K context | 204K | $0.225 | $0.90 | $0.045 | $0.2813 |
| Qwen | |||||
Qwen3.8 Maxqwen3.8-max1M context | 1M | $2.00 | $6.00 | $0.25 | $2.50 |
Qwen3.7 Plusqwen3.7-plus1M contextUp to 200K tokens | 1M | $0.30 | $1.20 | $0.03 | $0.375 |
| Over 200K tokens | $0.90 | $3.60 | $0.09 | $1.125 | |
Qwen3.7 Maxqwen3.7-max1M context | 1M | $1.875 | $5.625 | $0.375 | $2.344 |
Qwen3.6 Plusqwen3.6-plus1M contextUp to 200K tokens | 1M | $0.375 | $2.25 | $0.0375 | $0.4688 |
| Over 200K tokens | $1.50 | $4.50 | $0.15 | $1.875 | |
| Hy | |||||
Hy3hy3256K context | 256K | $0.105 | $0.435 | $0.02625 | — |
| LongCat | |||||
LongCat 2.0longcat-2.01M context | 1M | $0.225 | $0.90 | $0.0045 | — |
| Muse | |||||
Muse Spark 1.2muse-spark-1.21M context | 1M | $0.075 | $0.15 | $0.0015 | — |
| DeepSeek | |||||
DeepSeek V4 Prodeepseek-v4-pro1M context | 1M | $1.32 | $3.96 | $0.044 | — |
DeepSeek V4 Flash Vision Expdeepseek-v4-flash-vision-exp1M context | 1M | $0.44 | $1.32 | $0.014 | — |
DeepSeek V4 Flashdeepseek-v4-flash1M context | 1M | $0.286 | $0.858 | $0.0091 | — |
| Grok | |||||
Grok 4.6grok-4.6500K contextUp to 200K tokens | 500K | $2.00 | $6.00 | $0.50 | — |
| Over 200K tokens | $4.00 | $12.00 | $1.00 | — | |
| GPT | |||||
GPT 5.6 Lunagpt-5.6-luna1M contextUp to 200K tokens | 1M | $0.20 | $1.20 | $0.02 | $0.25 |
| Over 200K tokens | $0.40 | $1.80 | $0.04 | $0.50 | |
Cached input is billed at the cache-read rate automatically: the discount applies itself, with nothing to enable and no separate plan. A model with no cache-write rate does not charge for writing to the cache at all.
DeepSeek rates are flat around the clock: no peak or off-peak window, and no request costs more or less for the time it was sent.
Rates are shown rounded to four significant digits; your balance is charged at the exact published rate, which never differs from the figure shown by more than a tenth of a percent.