This is a log you can follow to reproduce any of my tests. I've run several 26B-class models on my LAN box 235h, and the pitfalls are all here. I'm writing it down because re-researching them on every model or backend switch is pure waste.
Machine at a glance
| Item | Value |
|---|
| Host | 235h — Windows 11 企业版 LTSC 26200 |
| IP | 192.168.31.61(SSH 免密) |
| CPU | Intel 14 核 14 线程,Max 2.0GHz |
| GPU | Intel Arc 140T 核显(12GB,UMA 共享) |
| RAM | 23.5 GB |
| iGPU shared VRAM | 默认 57%(约 13.4GB)→ 调优后 87-90%(约 19.7-21.2 GiB) |
First pitfallThe default 57% share can't hold a 16 GB 26B model —
Vulkan / SYCL≥24 / OpenVINO all fail with OutOfDeviceMemory or res=-3. You must raise
SystemPartitionCommitLimitPercentage first (see
iGPU shared VRAM tuning).
Model list: two routes
Route A: OpenVINO IR (official pre-converted)
| Property | Value |
|---|
| 模型 | gemma-4-26b-a4b-it-int4-ov |
| Type | OpenVINO IR,INT4(NNCF INT4_ASYM gs128,AWQ 精度) |
| Size | 20 文件 / 14.3 GB |
| Source | OpenVINO/gemma-4-26b-a4b-it-int4-ov(hf-mirror) |
The IR directory: openvino_language_model.bin is the bulk (13.0 GB), openvino_vision_embeddings_model.bin is the vision encoder (547 MB), plus tokenizer, detokenizer and text embeddings. Multimodal goes through VLMPipeline with image support.
Pitfall log: openvino_language_model.xml should be 5.1 MB. A truncated 118 KB copy floating around the LAN caused Error parsing element attribute at offset 118773 on load — re-download the correct file from hf-mirror.
Route B: GGUF (llama.cpp universal format)
| Model | Size | Params |
|---|
| gemma-4-26B-A4B-it-Q4_K_M.gguf | 16.02 GB | 26B.A4B |
| supergemma4-26b-uncensored-fast-v2-Q4_K_M.gguf | 16.02 GB | 25.23B |
Download sources
| Source | Note |
|---|
| hf-mirror(OpenVINO IR) | cdn 302 → us.aws.cdn.hf.co,实测 curl -L 约 26 MB/s |
| ggml-org(GGUF) | llama.cpp 官方原版转换 |
Route A results: OpenVINO GenAI native pipeline
The most important correction (2026-09-04): the correct backend is ATTENTION_BACKEND=PA, not SDPA. SDPA makes long-context decode collapse — about 2 t/s by 8k — while PA stays flat at ~14-15 t/s all the way to 16k. The early "you must use SDPA to bypass the continuous-batching assertion" conclusion was wrong.
import openvino_genai as ov_genai
MODEL_DIR = r"C:\path\ov-gemma4-26b"
# PA 后端:长上下文 decode 不掉速
# 32k 级长输入再加 KV_CACHE_PRECISION=u4
pipe = ov_genai.VLMPipeline(
MODEL_DIR, "GPU", **{"ATTENTION_BACKEND": "PA"}
)
res = pipe.generate(
"请用中文简述量子纠缠是什么,并给出一个应用例子。",
max_new_tokens=120,
)
print(res.texts[0])
Four API gotchas
- Use keyword arguments (
max_new_tokens=); don't pass a GenerationConfig object explicitly. temperature=0.0 is rejected (must be > 0 when sampling) — just use the default.- Read the output from
res.texts[0]. - You must use VLMPipeline — LLMPipeline fails to find
input_ids because the input is inputs_embeds.
GenerationConfig sampling defaults
| Field | Default | Note |
|---|
| do_sample | False | greedy, fastest; True samples and is slower |
| temperature | 1.0 | Sampling temperature |
| top_k / top_p | ∞ / 1.0 | Sampling switches |
Best config per route
| Use case | Best route | Key numbers |
|---|
| General inference | llama.cpp Vulkan 全量 offload | pp512 243.6 / tg128 21.4 t/s |
| Official IR weights | OpenVINO GenAI(PA 后端) | pp512 约 500 / decode 14-15 t/s |
| Long-context service | llama-server -ub 256 -c 65536 | 20K+ 稳定,prefill 69.3 t/s |
| 16K benchmarks | serve_ov.py(OV PA 后端) | 16k+500 输出约 67s 往返 |
Important MoE limitationFor MoE models like 26B.A4B, llama.cpp allocates weights per whole layer — it can't put active experts on GPU and inactive ones on CPU. So even with only 4B active parameters, the full 16 GB of weights must fit in VRAM. Few active parameters ≠ low memory use.
Serving: long-context parameters
llama-server.exe -m <model>.gguf \
-ngl 99 -t 8 -fa on \
-ub 256 -c 65536 \
--host 0.0.0.0 --port 8080 --cache-reuse 64
-c 65536 with -ub 256 is stable up to 30k; 20k prefill runs at 69.3 t/s.--cache-reuse 64 enables prefix caching so repeated prefixes skip most of the prefill.- Larger
-ub eats FA scratch memory — -c 65536 with -ub 512 gave me an iGPU driver-level OOM.
An architecture-level OpenVINO bug
The gemma4 architecture on llama.cpp's OpenVINO backend only works stateless and crashes at long context: E4B dies around pos 1024, 12B around 1536/2048, 26B likewise. The workaround is GGML_OPENVINO_STATEFUL_EXECUTION=0 with -fa 1.
If your model is from the gemma family, be careful with the OpenVINO backend — Vulkan is slower on prefill but far more stable at long context.
Reproduction quick reference
:: 测速(llama-bench)
cd C:\llama\vulkan
llama-bench.exe -m <gguf> -ngl 99 -t 8
:: 切到 OpenVINO 后端
set GGML_OPENVINO_DEVICE=GPU
cd C:\llama\openvino
llama-bench.exe -m <gguf> -ngl 99 -t 8
The finer test scripts (prefill sweeps, decode timing, memory monitoring, OV perf hints) are archived in the same directory — 40+ scripts total. The three I use most: ov_prefill.py, decode_timer.py, and monitor_mem.py.
Data noteAll figures are measured on my own box (235h / Arc 140T / llama.cpp b10621 / OpenVINO GenAI 2026.5 nightly) and vary by driver and config. For the full backend comparison see
llama.cpp four-backend benchmark, and for serving details see
llama-server serving.