同一个模型,各家服务商差多少?80 组实测数据横向对比

换模型的时候纠结半天选哪家服务商?其实同一个模型在不同服务商之间,速度差得比换模型还大。我抓了两个公开测速站的数据,对齐后拿 DeepSeek-V4-Flash 做了个横评。

站点定位币种特有指标
AI Ping清华系清程极智,国内云/代理渠道监测人民币P90 首字延迟、近 6 小时成功率
LMSpeed全球 100+ 卖家比价,五轮测速美元首字延迟中位数、近 30 天可用性

数据规模:AI Ping 覆盖 146 个模型、80 组服务商级实测;LMSpeed 覆盖 446 条目录记录。两站按模型对齐后共 54 个可比模型,其中 17 个两站都有数据。

DeepSeek-V4-Flash:价格一样,速度差 3 倍

这个模型官方价是输入 ¥1/M、输出 ¥2/M,缓存命中 ¥0.2/M。不同服务商报的价格完全一样,但性能差得很远:

服务商吞吐 (tok/s)首字延迟可靠性
并行智算云131.950.09 s99%
金山云星流123.093.62 s100%
百度智能云111.230.40 s70%
PPIO 派欧云———
DeepSeek-V4-Flash 在各家服务商的吞吐量对比,最高与最低差约三倍
吞吐对比(tok/s)
并行智算云131.95
金山云星流123.09
百度智能云111.23

吞吐最高和最低差约 3 倍(131.95 vs 40-60 区间的中位水平)。但首字延迟的差距更夸张:并行智算云 0.09 秒,金山云 3.62 秒,差 40 倍。对交互式场景来说,3.6 秒的首字延迟比吞吐慢 3 倍难受得多。

意外的发现价格不是性能指标。这几家报的价格一模一样,但速度和可靠性差出一个数量级。反过来,可靠性也不和速度正相关——百度智能云吞吐 111 t/s 排第三,可靠性只有 70%,是金山云 100% 的七折。

国内七家云厂商速度排名

这是 AI Ping 首页近 7 日的吞吐榜(TOP5)和首字延迟榜(TOP5),覆盖了主要的几家:

厂商吞吐 (tok/s)首字延迟 (s)
百度智能云114.99—
讯飞星辰106.17—
DeepSeek 官方100.770.18
PPIO 派欧云97.740.34
腾讯云 TokenHub93.050.55
火山方舟—1.26
快手万擎—1.30
吞吐 TOP5(tok/s)
百度智能云114.99
讯飞星辰106.17
DeepSeek100.77
PPIO97.74
腾讯 TokenHub93.05

两个榜单的交叉很有意思:DeepSeek 官方在延迟榜排第一(0.18 秒),吞吐榜只排第三。而百度智能云吞吐最高(114.99),延迟却没进前五。对延迟敏感的场景(比如实时对话、Agent 工具调用),DeepSeek 官方更靠谱。

选择服务商的四个维度

  • 首字延迟优先于吞吐。交互式场景用户等的是第一个字,不是每秒几个 token。3.6 秒的首字延迟比慢 3 倍的吞吐更致命。
  • 可靠性要和速度一起看。70% 的成功率意味着每三次请求失败一次,这种波动在生产环境里比慢几秒更难受。
  • 价格一致时,服务商选择就是纯性能题。DeepSeek-V4-Flash 各家报价一样,没什么理由为速度付溢价,也没什么理由不挑最快的。
  • 协议兼容性也要看。很多服务商同时支持 OpenAI Chat Completions 和 Anthropic Messages,迁移成本差别很大。

数据口径说明(别混着用)

  • 币种不同未换算:AI Ping 是人民币,LMSpeed 是美元,两者不能直接相减。
  • 测速口径不同:AI Ping 是每服务商一条 P90 首字延迟 + 平均吞吐;LMSpeed 是五轮测速取中位数。统计口径不一样,数字不宜并列。
  • 覆盖范围不同:AI Ping 只覆盖国内云和代理,LMSpeed 覆盖全球中转站。同一模型两站的"最快服务商"通常不是同一家。
  • LMSpeed 公开端点有限:五轮测速只对 20 个公开端点开放,登录才可见全部。
实用建议别只测一次就换服务商。这些数字是聚合的周期均值,单次调用受网络和负载影响很大。真要选的话,拿自己真实的 prompt 在候选服务商上各跑 20 次,比较 P50 和 P90 首字延迟,比看任何榜单都准。
数据说明抓取日期 2026-09-11,数据来自 AI Ping(aiping.cn)与 LMSpeed(lmspeed.net)公开页面,未登录抓取。两站数据每日更新,模型价格和服务商随时变动,本��数字仅代表该时点。价格与倍数的详细拆解见 缓存命中率 99.3% 账单拆解。

Debating which provider to pick for a model? Speed differences between providers of the same model are often larger than differences between models. I pulled data from two public benchmarking sites and lined it up, using DeepSeek-V4-Flash as the case study.

SiteFocusCurrencyUnique metric
AI Ping清华系清程极智,国内云/代理渠道监测人民币P90 首字延迟、近 6 小时成功率
LMSpeed全球 100+ 卖家比价,五轮测速美元首字延迟中位数、近 30 天可用性

Scale: AI Ping covers 146 models with 80 provider-level measurements; LMSpeed has 446 catalog entries. Aligned by model, 54 models are comparable, and 17 appear on both sites.

DeepSeek-V4-Flash: same price, 3x speed gap

Official pricing is ¥1/M input, ¥2/M output, ¥0.2/M cache hit. Providers quote identical prices while performance varies wildly:

ProviderThroughputTTFTReliability
并行智算云131.950.09 s99%
金山云星流123.093.62 s100%
百度智能云111.230.40 s70%
PPIO 派欧云———
DeepSeek-V4-Flash throughput across providers, with roughly a 3x spread between fastest and slowest
Throughput (tok/s)
并行智算云131.95
金山云星流123.09
百度智能云111.23

The fastest and slowest differ about 3x in throughput. But first-token latency is far more dramatic: 0.09s vs 3.62s — a 40x gap. For interactive work, a 3.6s TTFT hurts far more than 3x slower throughput.

Unexpected findingPrice is not a performance metric. Identical prices, but speed and reliability differ by an order of magnitude. And reliability doesn't correlate with speed — Baidu ranks third on throughput at 111 t/s yet only 70% reliable, versus Kingsoft Cloud at 100%.

Seven Chinese cloud vendors, ranked by speed

AI Ping's last-7-days throughput (top 5) and first-token latency (top 5) boards, covering the major vendors:

VendorThroughputTTFT
Baidu114.99—
iFlytek106.17—
DeepSeek (official)100.770.18
PPIO97.740.34
Tencent TokenHub93.050.55
Volcengine—1.26
Kuaishou—1.30
Throughput top 5 (tok/s)
百度智能云114.99
讯飞星辰106.17
DeepSeek100.77
PPIO97.74
腾讯 TokenHub93.05

The crossover is interesting: DeepSeek official is #1 on latency (0.18s) but only #3 on throughput. Baidu leads throughput (114.99) but didn't make the latency top 5. For latency-sensitive work (real-time chat, Agent tool calls), DeepSeek official is the safer pick.

Four dimensions when choosing a provider

  • TTFT beats throughput. Users wait for the first word, not tokens per second. A 3.6s TTFT hurts more than 3x slower throughput.
  • Check reliability alongside speed. 70% success means one in three requests fails — that variability hurts more in production than a few extra seconds.
  • When prices match, provider choice is pure performance. Identical quotes mean no reason to pay more — and no reason not to take the fastest.
  • Check protocol support too. Many providers serve both OpenAI Chat Completions and Anthropic Messages, which matters a lot for migration cost.

Measurement caveats (don't mix them)

  • No currency conversion: AI Ping is CNY, LMSpeed is USD — don't subtract one from the other.
  • Different methodologies: AI Ping gives a P90 TTFT plus average throughput; LMSpeed takes the median of five runs. Not directly comparable.
  • Different coverage: AI Ping is China-domestic only, LMSpeed covers global relays. The "fastest provider" differs between sites.
  • LMSpeed's public sample is limited: five-round tests cover only 20 public endpoints; the rest require an account.
Practical adviceDon't switch providers after one test. These are aggregated periodic averages; a single call varies with network and load. If you're actually choosing, run your own real prompt 20 times against each candidate and compare P50 and P90 TTFT — that's more accurate than any leaderboard.
Data noteFetched 2026-09-11 from public pages at AI Ping (aiping.cn) and LMSpeed (lmspeed.net), no login. Both update daily and prices/providers change constantly, so these figures reflect that moment only. For the cost breakdown see Breaking down a 99.3% cache-hit bill.
返回文章列表