KV 缓存到底吃多少显存?26B 模型 f16 vs u8 逐档实测

在核显上跑 26B 模型,显存是唯一的敌人。我的机器(235h)只有 23.5 GB 内存,核显共享显存预算 90% 约 21.2 GB,而模型静态占用就有 15.3 GB——留给 KV 缓存的只有 5 GB 出头。那 KV 到底吃多少?

我拆成两个变量扫:KV 精度(f16 / u8)× 输入长度,输出固定 100 token,每档 3 次取中位数。这是完整数据。

先看理论值

这个模型 30 层、8 个 KV 头、head_dim 256。每 token 的 KV = 2(K+V) × 8 × 256 × 30 × 位宽:

精度每 token30k 上下文
f16约 240 KB/token约 7.2 GB
u8约 120 KB/token约 3.6 GB

按理论算,30k 上下文用 u8 应该只要 3.6 GB,加上 15.3 GB 模型总共 18.9 GB,22 GB 预算内完全放得下。但实际呢?

实测:每一档的显存峰值

输入长度f16 峰值u8 峰值差值
54715.6515.370.28
1,06315.5515.310.24
2,05716.0215.590.43
4,03416.9116.220.69
8,04018.3017.580.72
8K 输入时 f16 与 u8 两种 KV 精度的显存峰值对比图

单位是 GB。模型静态占用 15.3 GB 是基线,所以 547 token 那档的 15.37 几乎就是纯模型权重。

8K 输入时显存对比(GB)
f1618.30
u817.58

反常识的发现:u8 只省 17%

从实测数据反推增长斜率:f16 约 0.35 KB/token,u8 约 0.29 KB/token。也就是说 u8 相比 f16 只降了约 17%,远不到理论上的 50%。

为什么没到理论值?因为实测斜率 0.29 KB/token 比理论 u8 值(0.12 KB/token)本身就高了一倍多。也就是说,显存里除了纯 KV 张量,还有 OpenVINO 运行时自己的临时缓冲、注意力工作区和量化元数据在占位置。
这意味着什么别指望 KV 量化救长上下文。如果你的显存卡在边缘(比如 8K 能跑、16K 就 OOM),把 f16 换成 u8 大概率救不了你——因为省下的 0.7 GB 往往还不够 8K 之后每多 4K 涨的量。得从别的地方找空间。

大场景验证:30k / 32k

小场景(≤8k)两边都轻松,测不出真相。所以我又压到 30k 和 32k:

结果是 32k 稳定跑通,没有触发 CL_OUT_OF_RESOURCES。这比我按小场景斜率外推的结果要好——小场景每 token 涨 0.29 KB,外推 32k 要多 9.3 GB,加上 15.3 GB 基线就是 24.6 GB,已经超过 21.2 GB 预算了,按理应该炸。

也就是说,OpenVINO 在长上下文时并不按小场景的斜率线性增长 KV。可能是分块处理(chunked prefill)导致工作区复用,也可能有更聪明的注意力实现。这一点官方文档没写清楚,我是实测出来的。

方法论:怎么测显存

UMA 架构下监测系统内存就是监测显存

这是核显最重要的一点:共享显存 = 系统内存。所以我直接看系统 free,不用去找什么 GPU 专用工具。我写了个 monitor_mem.py 常驻,每 5 秒写一次 mem_monitor.csv,调试时非常有用。

怎么控制显存的三个杠杆

  • 共享显存预算:注册表 SystemPartitionCommitLimitPercentage。这是最大的杠杆,默认 57% 调到 87-90%,预算从 13.4 GB 到 19.7-21.2 GB。
  • KV 精度:f16 → u8 有用但有限(实测省 17%),只在临界情况下能救一命。
  • 模型量化:INT4 IR vs INT8,权重本身是显存大头(15.3 GB),这一层的收益比 KV 量化大得多。

实践建议

  1. 先用小场景摸规律,再用大场景验证。8k 测出来的斜率不能用来外推 32k——我这次实测证明外推是错的。
  2. 别先纠结 KV 量化,先调共享显存预算。调 10% 预算带来的空间,比把 f16 换 u8 大一个量级。
  3. 贴着预算上限设计场景。算的时候留 10-15% 余量,别刚好卡死——运行时的临时缓冲不算在你算的 KV 里。
  4. 要稳过 32k,u4 值得试。我在 32k 级别用 KV_CACHE_PRECISION=u4 跑通了,虽然精度有损,但配合 PA 后端速度不掉。
数据说明本机实测(235h / Arc 140T / 23.5 GB UMA / OpenVINO GenAI 2026.5 nightly / gemma-4-26b int4 IR,PA 后端),输出固定 100 token,每档 3 次取中位数。显存数值含 OpenVINO 运行时开销,非纯 KV 大小。共享显存调优方法见 核显共享显存调优。

Running 26B on an iGPU means fighting for VRAM. My box (235h) has 23.5 GB RAM and a 90% shared-VRAM budget (≈21.2 GB), but the model itself takes 15.3 GB — leaving barely 5 GB for the KV cache. So how much does KV actually cost?

I varied two things independently: KV precision (f16 / u8) × input length, output fixed at 100 tokens, median of 3 runs. Here's the full data.

Theoretical baseline first

The model has 30 layers, 8 KV heads, head_dim 256. Per token: KV = 2(K+V) × 8 × 256 × 30 × bitwidth:

PrecisionPer tokenAt 30k context
f16约 240 KB/token约 7.2 GB
u8约 120 KB/token约 3.6 GB

In theory 30k at u8 needs 3.6 GB, so 15.3 + 3.6 = 18.9 GB total — comfortably inside the 22 GB budget. In practice?

Measured: peak VRAM at every tier

Input lengthf16 peaku8 peakDelta
54715.6515.370.28
1,06315.5515.310.24
2,05716.0215.590.43
4,03416.9116.220.69
8,04018.3017.580.72
VRAM peak comparison between f16 and u8 KV precision at 8K input

Units are GB. The 15.3 GB static model load is the baseline, so the 15.37 at 547 tokens is essentially pure weights.

VRAM at 8K input (GB)
f1618.30
u817.58

Counter-intuitive finding: u8 saves only 17%

Backing out the growth slope from the data: ≈0.35 KB/token for f16, ≈0.29 KB/token for u8 — so u8 saves roughly 17%, far short of the theoretical 50%.

Why the gap? The measured 0.29 KB/token is already more than double the theoretical u8 value (0.12 KB/token). Besides the raw KV tensors, OpenVINO's runtime scratch buffers, attention workspaces and quantization metadata all occupy VRAM.
What this meansDon't rely on KV quantization to save long context. If you're right at the edge (8K works, 16K OOMs), switching f16 → u8 probably won't save you — the 0.7 GB it frees is usually less than what the next 4K costs. Look elsewhere for headroom.

Large-scale verification: 30k / 32k

Both precisions are comfortable at small sizes, so the small sweep tells you nothing decisive. I pushed to 30k and 32k:

It ran 32k stably without CL_OUT_OF_RESOURCES — better than extrapolating from the small sweep. Extrapolating 0.29 KB/token to 32k predicts +9.3 GB, i.e. 24.6 GB on top of a 15.3 GB baseline, past the 21.2 GB budget. It should have failed.

In other words, OpenVINO does not grow KV linearly at long context the way the small sweep suggests. Possibly chunked prefill reuses workspaces, or the attention implementation is smarter than assumed. The docs don't say; I found this empirically.

Methodology: measuring VRAM correctly

On UMA, system memory is VRAM

The key fact about iGPUs: shared VRAM is system RAM. So I just watch system free instead of hunting for GPU-specific tools. My monitor_mem.py runs continuously, writing mem_monitor.csv every 5s — invaluable while debugging.

Three levers for controlling VRAM

  • Shared VRAM budget: the registry's SystemPartitionCommitLimitPercentage. The biggest lever — going from 57% to 87-90% moves the budget from 13.4 GB to 19.7–21.2 GB.
  • KV precision: f16 → u8 helps but modestly (17% measured), only enough to save you at the very edge.
  • Model quantization: weights are the bulk (15.3 GB), so gains here dwarf what KV quantization buys.

Practical recommendations

  1. Use small sweeps to find patterns, then verify large. An 8k slope cannot be extrapolated to 32k — my measurements prove that's wrong.
  2. Fix the shared-VRAM budget before touching KV precision. Ten percent more budget dwarfs the win from u8.
  3. Design with headroom. Leave 10-15% margin — runtime scratch buffers aren't in your KV calculation.
  4. For reliable 32k, try u4. I ran 32k with KV_CACHE_PRECISION=u4 — lossy on quality, but with the PA backend there's no speed penalty.
Data noteMeasured on my box (235h / Arc 140T / 23.5 GB UMA / OpenVINO GenAI 2026.5 nightly / gemma-4-26b int4 IR, PA backend), output fixed at 100 tokens, median of 3 runs. VRAM figures include OpenVINO runtime overhead, not raw KV size. For the budget tuning method see iGPU shared VRAM tuning.
返回文章列表