Pricing and Billing

What you are charged for, how caching lowers your cost, and how to check any single call against your bill.

1. What you pay for

Pricing is per token, quoted per 1 million tokens, and charged in CNY (¥). There is no subscription and no monthly minimum — you buy credits, and each call deducts what it actually used. Credits never expire.

Every call is billed across up to four separate token types, each with its own rate. This matters more than it sounds: output is typically 5–6× the input rate on the same model, so a rough estimate based on input alone will understate your bill.

Token typeTypical rateWhat it is
InputBaselineThe prompt you send, minus anything served from cache.
Output5–6× input on most modelsWhat the model generates. Usually the largest line on your bill — cap it with max_tokens if cost matters more than length. A few models price output far lower, so check the model you actually use.
Cache read~10% of input on most modelsPrompt content the provider already had cached from an earlier call. Heavily discounted — but not on every model: on a few, cached input costs the same as or more than regular input, so caching is not automatically a saving.
Cache writeVaries by modelContent being placed into the cache so later calls can reuse it. On Claude models this runs above the input rate; on several others it is below. Per-model rates are on the pricing page.

One API key reaches every model in your key's scope — you do not need a separate key per model. All of your keys draw on the same account balance.

2. All model prices

Every rate below is what your account is actually charged — these come from the same price records the billing system reads, so nothing here can drift out of sync with your bill. Prices are shown in USD, converted from CNY at 6.72 CNY/USD.

Models available

53

Providers

14

Lowest input price

$0.0387 / 1M

Exchange rate

6.72 CNY/USD

Showing 53 of 53 models

JZS Token1 model

jzs-max-3.0

JZS Token

Input / 1M
$0.0619
Output / 1M
$0.3095
Cache read
$0.0062
Cache write
$0.0619
1MUnderstands imagesExclusive

Alibaba / Qwen7 models

qwen3.6-flash

Alibaba / Qwen

Input / 1M
$0.1548

38% below the vendor's list price

Output / 1M
$0.7738
Cache read
$0.0155
Cache write
$0.1548
—Image support unknown
qwen3.8-flash

Alibaba / Qwen

Input / 1M
$0.1548
Output / 1M
$0.7738
Cache read
$0.0155
Cache write
$0.1548
—Image support unknown
qwen3.6-plus

Alibaba / Qwen

Input / 1M
$0.2321

53% below the vendor's list price

Output / 1M
$1.16
Cache read
$0.0232
Cache write
$0.2321
1MUnderstands images
qwen3.7-plus

Alibaba / Qwen

Input / 1M
$0.3095
Output / 1M
$1.55
Cache read
$0.0310
Cache write
$0.3095
1MUnderstands images
qwen3.8-omni-flash

Alibaba / Qwen

Input / 1M
$0.3869
Output / 1M
$1.93
Cache read
$0.0387
Cache write
$0.3869
—Image support unknown
qwen3.7-max

Alibaba / Qwen

Input / 1M
$0.7738

69% below the vendor's list price

Output / 1M
$2.32
Cache read
$0.0967
Cache write
$0.9673
1MText input only
qwen3.8-max

Alibaba / Qwen

Input / 1M
$1.55

22% below the vendor's list price

Output / 1M
$4.64
Cache read
$0.1935
Cache write
$1.93
1MUnderstands images

Anthropic7 models

Input / 1M
$0.0774

92% below the vendor's list price

Output / 1M
$0.3869
Cache read
$0.0077
Cache write
$0.1161
256KUnderstands images
Input / 1M
$0.7738

61% below the vendor's list price

Output / 1M
$3.87
Cache read
$0.0774
Cache write
$0.9673
1MUnderstands images
Input / 1M
$1.16
Output / 1M
$5.80
Cache read
$0.1161
Cache write
$1.45
1MUnderstands images
Input / 1M
$1.16
Output / 1M
$5.80
Cache read
$0.1161
Cache write
$1.45
1MUnderstands images
Input / 1M
$1.16
Output / 1M
$5.80
Cache read
$0.1161
Cache write
$1.45
1MUnderstands images
claude-opus-5

Anthropic

Input / 1M
$1.16

76% below the vendor's list price

Output / 1M
$5.80
Cache read
$0.1161
Cache write
$1.45
1MUnderstands images
Input / 1M
$1.55

69% below the vendor's list price

Output / 1M
$7.74
Cache read
$0.0155
Cache write
$1.93
—Image support unknown

ByteDance / Doubao4 models

doubao-seed-2.0-lite

ByteDance / Doubao

Input / 1M
$0.0774

8% below the vendor's list price

Output / 1M
$0.3869
Cache read
$0.0077
Cache write
$0.0387
—Image support unknown
doubao-seed-2.0-code

ByteDance / Doubao

Input / 1M
$0.2321

48% below the vendor's list price

Output / 1M
$1.16
Cache read
$0.0232
Cache write
$0.2321
200KUnderstands images
doubao-seed-2.0-pro

ByteDance / Doubao

Input / 1M
$0.2321

48% below the vendor's list price

Output / 1M
$1.16
Cache read
$0.0232
Cache write
$0.2321
128KUnderstands images
doubao-seed-2.1-turbo

ByteDance / Doubao

Input / 1M
$0.2321

45% below the vendor's list price

Output / 1M
$1.16
Cache read
$0.0232
Cache write
$0.2321
—Image support unknown

DeepSeek3 models

Input / 1M
$0.1488
Output / 1M
$0.7440
Cache read
$0.0149
Cache write
$0.1488
1MText input only⏱ peak 09:00–12:00 ×2, 14:00–18:00 ×2 (Asia/Shanghai)
Input / 1M
$0.1488
Output / 1M
$0.7440
Cache read
$0.0149
Cache write
$0.1488
—Image support unknown⏱ peak 09:00–12:00 ×2, 14:00–18:00 ×2 (Asia/Shanghai)
Input / 1M
$0.4464
Output / 1M
$2.23
Cache read
$0.0446
Cache write
$0.4464
1MText input only⏱ peak 09:00–12:00 ×2, 14:00–18:00 ×2 (Asia/Shanghai)

Google / Gemini3 models

gemini-3-flash

Google / Gemini

Input / 1M
$0.2321
Output / 1M
$1.16
Cache read
$0.0232
Cache write
$0.1161
—Image support unknown
gemini-3.5-flash

Google / Gemini

Input / 1M
$0.4643
Output / 1M
$2.32
Cache read
$0.0464
Cache write
$0.2321
—Image support unknown
gemini-3.1-pro

Google / Gemini

Input / 1M
$0.7738
Output / 1M
$3.87
Cache read
$0.0774
Cache write
$0.3869
—Image support unknown

Meituan / LongCat1 model

LongCat-2.0

Meituan / LongCat

Input / 1M
$0.0774

74% below the vendor's list price

Output / 1M
$0.3095
Cache read
$0.0077
Cache write
$0.0774
1MText input only

MiniMax4 models

Input / 1M
$0.0387

83% below the vendor's list price

Output / 1M
$0.1935
Cache read
$0.0039
Cache write
$0.0193
—Image support unknown
Input / 1M
$0.0774

83% below the vendor's list price

Output / 1M
$0.3869
Cache read
$0.0077
Cache write
$0.0387
—Image support unknown
MiniMax-M3

MiniMax

Input / 1M
$0.0774

67% below the vendor's list price

Output / 1M
$0.3869
Cache read
$0.0077
Cache write
$0.0193
1MUnderstands images
Input / 1M
$0.1548
Output / 1M
$0.7738
Cache read
$0.0155
Cache write
$0.0774
—Image support unknown

Moonshot / Kimi3 models

kimi-k2.6

Moonshot / Kimi

Input / 1M
$0.1548

80% below the vendor's list price

Output / 1M
$0.7738
Cache read
$0.0155
Cache write
$0.1548
—Image support unknown
kimi-k2.7

Moonshot / Kimi

Input / 1M
$0.4643
Output / 1M
$2.32
Cache read
$0.0464
Cache write
$0.4643
—Image support unknown
kimi-k3

Moonshot / Kimi

Input / 1M
$1.93

35% below the vendor's list price

Output / 1M
$9.67
Cache read
$0.1935
Cache write
$1.93
1MUnderstands images

OpenAI9 models

Input / 1M
$0.0387
Output / 1M
$0.2321
Cache read
$0.0039
Cache write
$0.0484
—Image support unknown
Input / 1M
$0.0774
Output / 1M
$0.0464
Cache read
$0.0077
Cache write
$0.1161
258KUnderstands images
Input / 1M
$0.2321
Output / 1M
$1.39
Cache read
$0.0232
Cache write
$0.3482
258KUnderstands images
gpt-6-sol

OpenAI

Input / 1M
$0.2321
Output / 1M
$1.39
Cache read
$0.0232
Cache write
$0.2902
—Image support unknown
gpt-5.5

OpenAI

Input / 1M
$0.3869
Output / 1M
$2.32
Cache read
$0.0387
Cache write
$0.3869
258KUnderstands images
Input / 1M
$0.4643
Output / 1M
$2.79
Cache read
$0.0464
Cache write
$0.5804
258KUnderstands images
Input / 1M
$1.16
Output / 1M
$6.96
Cache read
$0.1161
Cache write
$1.74
—Image support unknown

$0.1935 per image

4K outputGenerates images

$0.0193 per image

1–2K outputGenerates images

StepFun1 model

Input / 1M
$0.0774

59% below the vendor's list price

Output / 1M
$0.3869
Cache read
$0.0077
Cache write
$0.0387
256KUnderstands images

xAI / Grok3 models

grok-4.5

xAI / Grok

Input / 1M
$0.3869
Output / 1M
$1.93
Cache read
$0.4836
Cache write
$0.3869
500KUnderstands images
grok-4.6

xAI / Grok

Input / 1M
$0.3869
Output / 1M
$1.93
Cache read
$0.4836
Cache write
$0.3869
—Image support unknown
grok-4.7

xAI / Grok

Input / 1M
$0.3869
Output / 1M
$1.93
Cache read
$0.0387
Cache write
$0.4836
—Image support unknown

Xiaomi / MiMo4 models

mimo-v2.5

Xiaomi / MiMo

Input / 1M
$0.1548
Output / 1M
$0.7738
Cache read
$0.0155
Cache write
$0.1548
1MUnderstands images
mimo-v2.6-flash

Xiaomi / MiMo

Input / 1M
$0.1548
Output / 1M
$0.7738
Cache read
$0.0155
Cache write
$0.1548
—Image support unknown
mimo-v2.5-pro

Xiaomi / MiMo

Input / 1M
$0.3095
Output / 1M
$1.55
Cache read
$0.0310
Cache write
$0.3095
1MText input only
mimo-v2.6-pro

Xiaomi / MiMo

Input / 1M
$0.3095
Output / 1M
$1.55
Cache read
$0.0310
Cache write
$0.3095
—Image support unknown

Zhipu / Z.ai3 models

glm-5.3-flash

Zhipu / Z.ai

Input / 1M
$0.1548
Output / 1M
$0.7738
Cache read
$0.0155
Cache write
$0.1548
—Image support unknown
glm-5.2

Zhipu / Z.ai

Input / 1M
$0.6190

29% below the vendor's list price

Output / 1M
$3.10
Cache read
$0.0619
Cache write
$0.6190
1MUnderstands images
glm-5.3

Zhipu / Z.ai

Input / 1M
$0.6190

29% below the vendor's list price

Output / 1M
$3.10
Cache read
$0.0619
Cache write
$0.6190
—Image support unknown

Last updated Sep 24, 2026, 10:20 AM UTC. The pricing page shows the same catalogue with model comparisons and copyable model IDs.

3. Caching, and how to actually benefit from it

When consecutive requests share a long identical prefix, the provider can reuse the work it already did on that prefix instead of reprocessing it. You are charged the discounted cache-read rate for that portion. Long conversations, follow-up questions on the same document, code completion, and agent loops all hit this case naturally.

The practical rule: keep the unchanging part of your prompt at the front, and put what varies at the end. Caching matches on a shared prefix, so a system prompt or document that sits before your question stays cacheable across turns. Move a timestamp or a request ID to the top and you invalidate the prefix on every single call — you then pay full input rate for the whole thing, plus a cache write.

Cache writes are not free, so a one-off call that will never be followed up gains nothing from caching. The benefit shows up from the second call onward, which is exactly the pattern agents and multi-turn chat produce.

4. Image models are billed per image

Image generation is not billed per token. Each image is a fixed price regardless of prompt length, and the rate depends on the output resolution. Current per-image prices are on the pricing page.

Image calls go to POST /v1/images/generations, not the chat endpoint. In your usage records they show a cost with zero tokens — that is expected, since tokens play no part in what you were charged.

5. Checking a call against your bill

The usage lookup page takes an API key and shows your balance, per-model totals, and recent calls. The four token types appear as separate columns, so you can multiply each one by its rate on the pricing page and land on the exact figure that was deducted.

The token counts come from the provider's usage block on the response — the same numbers your own client receives, so you can verify them against your logs without asking us.

Failed calls are listed too, with the reason in the Details column and a cost of zero. A call that errored is never billed. If the failure came from the upstream provider rather than your request, that column says so — you do not need to go debugging your own code first.

6. When prices change

Model prices track what the providers charge, and those move occasionally. Our price records sync automatically and the pricing page shows when they were last updated. Each call is billed at the rate in effect at that moment, and your usage records keep the resulting cost, so a later price change never restates what you already paid.

A few models are priced by time of day, meaning a call during the provider's peak window costs more than the same call off-peak. Where that applies, the pricing page notes it on the model. If your workload is flexible, shifting it out of the peak window is the simplest saving available.

Questions about a specific charge

Send us the call's timestamp and the model, plus the masked form of your API key — never a full working key. See the contact page for how to reach us.