Running 26B on an iGPU means fighting for VRAM. My box (235h) has 23.5 GB RAM and a 90% shared-VRAM budget (≈21.2 GB), but the model itself takes 15.3 GB — leaving barely 5 GB for the KV cache. So how much does KV actually cost?
I varied two things independently: KV precision (f16 / u8) × input length, output fixed at 100 tokens, median of 3 runs. Here's the full data.
Theoretical baseline first
The model has 30 layers, 8 KV heads, head_dim 256. Per token: KV = 2(K+V) × 8 × 256 × 30 × bitwidth:
| Precision | Per token | At 30k context |
|---|
| f16 | 约 240 KB/token | 约 7.2 GB |
| u8 | 约 120 KB/token | 约 3.6 GB |
In theory 30k at u8 needs 3.6 GB, so 15.3 + 3.6 = 18.9 GB total — comfortably inside the 22 GB budget. In practice?
Measured: peak VRAM at every tier
| Input length | f16 peak | u8 peak | Delta |
|---|
| 547 | 15.65 | 15.37 | 0.28 |
| 1,063 | 15.55 | 15.31 | 0.24 |
| 2,057 | 16.02 | 15.59 | 0.43 |
| 4,034 | 16.91 | 16.22 | 0.69 |
| 8,040 | 18.30 | 17.58 | 0.72 |
Units are GB. The 15.3 GB static model load is the baseline, so the 15.37 at 547 tokens is essentially pure weights.
Counter-intuitive finding: u8 saves only 17%
Backing out the growth slope from the data: ≈0.35 KB/token for f16, ≈0.29 KB/token for u8 — so u8 saves roughly 17%, far short of the theoretical 50%.
Why the gap? The measured 0.29 KB/token is already more than double the theoretical u8 value (0.12 KB/token). Besides the raw KV tensors, OpenVINO's runtime scratch buffers, attention workspaces and quantization metadata all occupy VRAM.
What this meansDon't rely on KV quantization to save long context. If you're right at the edge (8K works, 16K OOMs), switching f16 → u8 probably won't save you — the 0.7 GB it frees is usually less than what the next 4K costs. Look elsewhere for headroom.
Large-scale verification: 30k / 32k
Both precisions are comfortable at small sizes, so the small sweep tells you nothing decisive. I pushed to 30k and 32k:
It ran 32k stably without CL_OUT_OF_RESOURCES — better than extrapolating from the small sweep. Extrapolating 0.29 KB/token to 32k predicts +9.3 GB, i.e. 24.6 GB on top of a 15.3 GB baseline, past the 21.2 GB budget. It should have failed.
In other words, OpenVINO does not grow KV linearly at long context the way the small sweep suggests. Possibly chunked prefill reuses workspaces, or the attention implementation is smarter than assumed. The docs don't say; I found this empirically.
Methodology: measuring VRAM correctly
On UMA, system memory is VRAM
The key fact about iGPUs: shared VRAM is system RAM. So I just watch system free instead of hunting for GPU-specific tools. My monitor_mem.py runs continuously, writing mem_monitor.csv every 5s — invaluable while debugging.
Three levers for controlling VRAM
- Shared VRAM budget: the registry's
SystemPartitionCommitLimitPercentage. The biggest lever — going from 57% to 87-90% moves the budget from 13.4 GB to 19.7–21.2 GB. - KV precision: f16 → u8 helps but modestly (17% measured), only enough to save you at the very edge.
- Model quantization: weights are the bulk (15.3 GB), so gains here dwarf what KV quantization buys.
Practical recommendations
- Use small sweeps to find patterns, then verify large. An 8k slope cannot be extrapolated to 32k — my measurements prove that's wrong.
- Fix the shared-VRAM budget before touching KV precision. Ten percent more budget dwarfs the win from u8.
- Design with headroom. Leave 10-15% margin — runtime scratch buffers aren't in your KV calculation.
- For reliable 32k, try u4. I ran 32k with
KV_CACHE_PRECISION=u4 — lossy on quality, but with the PA backend there's no speed penalty.
Data noteMeasured on my box (235h / Arc 140T / 23.5 GB UMA / OpenVINO GenAI 2026.5 nightly / gemma-4-26b int4 IR, PA backend), output fixed at 100 tokens, median of 3 runs. VRAM figures include OpenVINO runtime overhead, not raw KV size. For the budget tuning method see
iGPU shared VRAM tuning.