Debating which provider to pick for a model? Speed differences between providers of the same model are often larger than differences between models. I pulled data from two public benchmarking sites and lined it up, using DeepSeek-V4-Flash as the case study.
| Site | Focus | Currency | Unique metric |
|---|
| AI Ping | 清华系清程极智,国内云/代理渠道监测 | 人民币 | P90 首字延迟、近 6 小时成功率 |
| LMSpeed | 全球 100+ 卖家比价,五轮测速 | 美元 | 首字延迟中位数、近 30 天可用性 |
Scale: AI Ping covers 146 models with 80 provider-level measurements; LMSpeed has 446 catalog entries. Aligned by model, 54 models are comparable, and 17 appear on both sites.
DeepSeek-V4-Flash: same price, 3x speed gap
Official pricing is ¥1/M input, ¥2/M output, ¥0.2/M cache hit. Providers quote identical prices while performance varies wildly:
| Provider | Throughput | TTFT | Reliability |
|---|
| 并行智算云 | 131.95 | 0.09 s | 99% |
| 金山云星流 | 123.09 | 3.62 s | 100% |
| 百度智能云 | 111.23 | 0.40 s | 70% |
| PPIO 派欧云 | — | — | — |
Throughput (tok/s)
并行智算云131.95
金山云星流123.09
百度智能云111.23
The fastest and slowest differ about 3x in throughput. But first-token latency is far more dramatic: 0.09s vs 3.62s — a 40x gap. For interactive work, a 3.6s TTFT hurts far more than 3x slower throughput.
Unexpected findingPrice is not a performance metric. Identical prices, but speed and reliability differ by an order of magnitude. And reliability doesn't correlate with speed — Baidu ranks third on throughput at 111 t/s yet only 70% reliable, versus Kingsoft Cloud at 100%.
Seven Chinese cloud vendors, ranked by speed
AI Ping's last-7-days throughput (top 5) and first-token latency (top 5) boards, covering the major vendors:
| Vendor | Throughput | TTFT |
|---|
| Baidu | 114.99 | — |
| iFlytek | 106.17 | — |
| DeepSeek (official) | 100.77 | 0.18 |
| PPIO | 97.74 | 0.34 |
| Tencent TokenHub | 93.05 | 0.55 |
| Volcengine | — | 1.26 |
| Kuaishou | — | 1.30 |
Throughput top 5 (tok/s)
百度智能云114.99
讯飞星辰106.17
DeepSeek100.77
PPIO97.74
腾讯 TokenHub93.05
The crossover is interesting: DeepSeek official is #1 on latency (0.18s) but only #3 on throughput. Baidu leads throughput (114.99) but didn't make the latency top 5. For latency-sensitive work (real-time chat, Agent tool calls), DeepSeek official is the safer pick.
Four dimensions when choosing a provider
- TTFT beats throughput. Users wait for the first word, not tokens per second. A 3.6s TTFT hurts more than 3x slower throughput.
- Check reliability alongside speed. 70% success means one in three requests fails — that variability hurts more in production than a few extra seconds.
- When prices match, provider choice is pure performance. Identical quotes mean no reason to pay more — and no reason not to take the fastest.
- Check protocol support too. Many providers serve both OpenAI Chat Completions and Anthropic Messages, which matters a lot for migration cost.
Measurement caveats (don't mix them)
- No currency conversion: AI Ping is CNY, LMSpeed is USD — don't subtract one from the other.
- Different methodologies: AI Ping gives a P90 TTFT plus average throughput; LMSpeed takes the median of five runs. Not directly comparable.
- Different coverage: AI Ping is China-domestic only, LMSpeed covers global relays. The "fastest provider" differs between sites.
- LMSpeed's public sample is limited: five-round tests cover only 20 public endpoints; the rest require an account.
Practical adviceDon't switch providers after one test. These are aggregated periodic averages; a single call varies with network and load. If you're actually choosing, run your own real prompt 20 times against each candidate and compare P50 and P90 TTFT — that's more accurate than any leaderboard.
Data noteFetched 2026-09-11 from public pages at AI Ping (aiping.cn) and LMSpeed (lmspeed.net), no login. Both update daily and prices/providers change constantly, so these figures reflect that moment only. For the cost breakdown see
Breaking down a 99.3% cache-hit bill.