上一篇按 Agent 场景把花费折算完了,这篇看「裸价」:官方刊例的输入单价、输出单价、缓存命中单价,三个榜全部从高到低排,共 109 款全量(缓存榜 70 款,只含有缓存报价的模型)。单价榜不掺任何模拟,就是官网标价,方便你按自己的真实用量去套。

价格为 2026 年 8 月官网刊例(元/百万 tokens),阶梯/分段计费取基础档,峰谷计价分别列出。

输入单价榜(节选 Top 20)

输入单价直接决定 Agent 的成本大头。破 10 元的几乎都和语音/多模态有关:

#模型(厂商)输入输出缓存
1Baichuan4(2024 定价)(百川)¥100100
2tencent-yuanqi(腾讯)¥100100
3step-1o-audio(语音)(阶跃)¥25605
4qwen3.8-max-prime(阿里)¥24722.4
5Baichuan3-Turbo-128k(百川)¥2424
6Baichuan2-53B 峰时(百川)¥2020
7kimi-k3(Moonshot)¥201002
8Baichuan4-Turbo(百川)¥1515
9kimi-k2.7-code 高速版(Moonshot)¥13542.6
10qwen3.8-max(阿里)¥12361.2
11Baichuan3-Turbo(百川)¥1212
12Baichuan-M3 / M2-Plus(百川)¥1030
13stepaudio 2.5 / 2 系列(语音)(阶跃)¥1025~1052
14GLM-4-AirX(智谱)¥1010
15deepseek-v4-pro-peak(DeepSeek)¥9270.3
16GLM-5.3 / GLM-5.2(智谱)¥8282
17Baichuan2-Turbo(百川)¥88
18kimi-k2.7-code(Moonshot)¥6.5271.3
19doubao-seed-2.1-pro 基础档(字节)¥6301.2
20qwen3.7-max 5 折(阿里)¥6180.6

输出单价榜(节选 Top 15)

输出贵的全是"长输出"型选手:语音转录、代码模型、旗舰推理。

#模型(厂商)输出
1Baichuan4 / tencent-yuanqi(2024 定价/平台价)¥100
2kimi-k3(Moonshot)¥100
3step-audio-r1.5(语音)(阶跃)¥105
4step-1o-audio(语音)(阶跃)¥60
5qwen3.8-max-prime(阿里)¥72
6kimi-k2.7-code 高速版(Moonshot)¥54
7Baichuan-M3 / M2-Plus(百川)¥30
8doubao-seed-2.1-pro / evolving(字节)¥30
9qwen3.8-max(阿里)¥36
10kimi-k2.7-code / k2.6(Moonshot)¥27
11deepseek-v4-pro-peak(DeepSeek)¥27
12GLM-5.3 / GLM-5.2(智谱)¥28
13Baichuan2-53B 峰时(百川)¥20
14minimax-m3-priority-gt512k 5 折(MiniMax)¥25.2
15stepaudio 系列(语音)(阶跃)¥70

缓存命中价榜(70 款有缓存报价,节选 Top 15)

缓存命中价是 Agent 场景的命门——它和输入原价的比值决定了重复上下文能省多少:

#模型(厂商)缓存命中占输入原价
1step-1o-audio(语音)(阶跃)¥520%
2qwen3.8-max-prime(阿里)¥2.410%
3kimi-k2.7-code 高速版(Moonshot)¥2.620%
4kimi-k3(Moonshot)¥210%
5stepaudio 系列(语音)(阶跃)¥220%
6GLM-5.3 / GLM-5.2(智谱)¥225%
7GLM-5.1(智谱)¥1.322%
8kimi-k2.7-code(Moonshot)¥1.320%
9qwen3.8-max(阿里)¥1.210%
10doubao-seed-2.1-pro(字节)¥1.220%
11kimi-k2.6(Moonshot)¥1.117%
12GLM-5V-Turbo / GLM-5-Turbo(智谱)¥1.224%
13GLM-5(智谱)¥125%
14deepseek-v4-pro-peak(DeepSeek)¥0.33.3%
15hy4-preview(腾讯)¥0.35%

注意 DeepSeek 和小米 MiMo 这两个异类:缓存命中价只有原价的 1~3%(deepseek 0.3/9、mimo 0.025/3),重上下文场景下成本优势是数量级的。

三榜完整长图

输入单价榜(1/6~3/6 节选头部)

输入单价榜 1/6
输入单价榜 1/6
输入单价榜 2/6
2/6
输入单价榜 3/6
3/6

输出单价榜(节选头部)

输出单价榜 1/6
输出单价榜 1/6
输出单价榜 2/6
2/6

缓存命中单价榜(节选头部)

缓存命中单价榜 1/4
缓存命中榜 1/4
缓存命中单价榜 2/4
2/4

怎么用这三张表

  • 按自己的用量套:你的场景输入输出比是多少、有没有重复前缀,直接拿裸价乘就行,比看任何模拟榜都准
  • Agent 选型先看缓存比:缓存命中价 ÷ 输入价越小,长上下文 Agent 越省钱
  • 峰谷价能薅:DeepSeek 谷时(offpeak)输入只有峰时一半,跑批量任务把时间挪到谷时就是半价

同系列:第 1 篇 Agent 场景算账总榜 · 第 3 篇厂商内部分榜

免责声明:价格为 2026 年 8 月官网刊例(元/百万 tokens),手工整理可能不全或有错漏,限时折扣按当前标注价,波动频繁仅供参考。

Where the previous post converted prices into simulated agent workloads, this one shows the raw sticker prices: official input, output and cache-hit rates for 109 Chinese LLM APIs (cache board covers 70 models with cache pricing), ranked high to low. August 2026 list prices in ¥/M tokens; tiered pricing at base tier, peak/off-peak listed separately.

Input price board (top 20 of 109)

#Model (vendor)InOutCache
1Baichuan4 (2024 legacy)¥100100
3step-1o-audio (voice)¥25605
4qwen3.8-max-prime¥24722.4
7kimi-k3¥201002
10qwen3.8-max¥12361.2
15deepseek-v4-pro-peak¥9270.3
16GLM-5.3 / 5.2¥8282
19doubao-seed-2.1-pro¥6301.2

Everything above ¥10/M input is voice or multimodal — text agents rarely touch this zone.

Output price board

Long-output models dominate: voice transcription (step-audio-r1.5 at ¥105), reasoning flagships (kimi-k3 at ¥100), and coding models (kimi-k2.7-code at ¥54).

Cache-hit board (the agent lever)

The cache-hit-to-input ratio decides how much repeated context saves:

  • Typical: 10–25% of input price (Qwen 1.2/12 = 10%, GLM 2/8 = 25%, Kimi 2/20 = 10%)
  • Outliers: DeepSeek at 0.3/9 = 3.3% and Xiaomi MiMo at 0.025/3 ≈ 0.8% — order-of-magnitude savings for long-context agents

How to use these boards

  • Plug your real input/output ratio into the sticker prices — more accurate than any simulation
  • For agents, pick by cache ratio first
  • Off-peak rates (DeepSeek input halves off-peak) effectively give 50% off batch jobs

Series: part 1 — agent workload costs · part 3 — per-vendor breakdowns.

Disclaimer: official list prices (¥/M tokens), August 2026; hand-collected and volatile — verify on official pages.
返回文章列表