自建 OpenAI 兼容 API 实测:7 个必踩的坑和 3 个致命问题

现在谁都在本地跑大模型,把它包成 OpenAI 兼容接口(图灵先森、LiteLLM、各种 one-api)之后,上层应用一行代码都不用改就能接上。但"兼容"两个字,水很深——我逐条测完发现 12 个问题,其中 3 个会直接让业务崩掉。

被测对象:局域网小主机 235h 上跑的 OpenVINO GenAI 服务,模型 gemma-4-26B(INT4,含视觉编码器),/health 返回 {"status":"ok","max_input_tokens":32000,"default_max_new":20480}。以下所有结论都可以复现。

先说好消息:核心路径全通

项目结果
非流式 chat/completions正常,结构标准
流式(SSE)正常,末尾 [DONE]
多轮对话 messages正常(能记住名字)
system 角色正常
max_tokens 别名max_new_tokens 也生效
图片多模态正常,识别很准
中文 UTF-8全程正常

流式响应里每段都带 usage 块(OpenAI 标准需要 stream_options.include_usage 才有),还附带了 prompt_tps / decode_tps / prefill_ms / decode_ms 四个性能字段。这是超集,做性能监控可以直接用,不用自己埋点。

坑 1:并发即 500(最致命)

两个请求同时发出(各 12k token),结果是一个 200 一个 500。报错体是 OpenVINO 的内部错误:Generate cannot be called while ContinuousBatchingPipeline is already in running state.

根因:HTTP 多线程 + 管线单实例

  • HTTP 层用 ThreadingHTTPServer(多线程),能同时收并发。
  • 但底层 VLMPipeline.generate 是单个全局实例、非线程安全。
  • 两个请求真实重叠时,后到的直接把 generate() 撞在"管线正在运行"上 → 立刻 500,而不是排队。
怎么办串行调用完全没问题(逐个发、等上一个完成)——这正是它"看起来像单线程"的原因。但如果将来有多个客户端/脚本同时打,会出现随机 500,没有重试逻辑业务直接失败。解决办法:在服务前加一层串行化队列或锁,或者客户端强制串行 + 超时重试。不要指望 OpenVINO 层自己排队。

坑 2:max_tokens 传错类型直接掐断连接

传 "max_tokens": "abc"(非数字字符串),连接直接被掐断,RemoteDisconnected,没有任何错误体。客户端只会以为网络故障去重试。

再测 max_tokens=0 或 -1,直接 500,而且错误体泄露 OpenVINO 内部 C++ 错误和磁盘路径:C:\Jenkins\workspace\...pipeline_impl.cpp。

根因是源码里 int(...) 抛 ValueError 没被捕获,handler 线程异常退出。修复很简单:入参校验 + 统一错误响应。

坑 3-7:五个兼容性缺口

3. 缺 /v1/models 端点 → 部分客户端联不上

GET /v1/models 返回 404。很多 OpenAI 客户端和 dashboard 在联调前会先拉模型列表(client.models.list()),直接就联不上。被测服务只实现了 /v1/chat/completions 和 /health。

4. finish_reason 永远是 length

连"1+1 等于几"答"2"这种自然结束的情况,也返回 finish_reason: "length"(源码硬编码)。文档里写"固定 length"是有意设计,但对 OpenAI 客户端是坑:很多 SDK 用 finish_reason == "stop" 判断回答是否完整、是否走 tool-call 分支,会全部误判。

5. temperature 等参数静默忽略

temperature / top_p / presence_penalty / seed / n / model 全部静默忽略——传了不报错也不生效,因为源码根本没读这些字段。客户端传 temperature=0.8 以为在采样,实际服务端是 do_sample=False 固定 greedy。这种"传了等于没传"比直接报错更危险,因为用户会误以为参数生效了。

6. 图片只支持两种传法

http(s):// 图片链接直接 400(不支持 http(s) 图片链接),因为没实现下载。支持 base64 data URI 和服务端本机文件路径。

安全问题本机文件路径这种传法等于服务端任意文件读取。实测我把路径设成 C:\Windows\web\wallpaper\Windows\img0.jpg,它真的读出来并描述了 Win11 壁纸。内网自用问题不大,但这套服务一旦暴露公网就是漏洞。我的修复是按来源限制:base64 远程/本机都允许,本地路径仅限 localhost 调用。

7. markdown 传图会翻车

用 ![img](data:image/png;base64,...) 这种 markdown 写法,服务端不解析,整段 base64 被当普通文本喂进模型。实测输出了一堆完全虚构的"AABB CCDD…"占位图描述,白烧 6k+ token。OpenAI 客户端里少数用 markdown 传图的方式会直接踩中。

速度实测:符合小主机预期

指标实测
冷 prefill430~630 t/s
decode12~16 t/s
20k 输入首 token 等待~45 秒
20k 前缀缓存命中后1.19 秒
自建 OpenAI 兼容接口的测试结果汇总:7 项通过、5 个兼容性缺口、2 个致命问题

冷启动 2 万 token 的首 token 要等 45 秒——这个体感很关键。但前缀缓存一旦命中就完全不是回事了:

调用总耗时TTFT说明
第 1 次(冷)45.1s44.5s完整 prefill
第 2 次(同前缀不同问)1.19s0.98s命中缓存
第 3 次(完全重复)0.58s0.44s全缓存

20k 前缀冷→热提速约 38 倍。生效条件是前缀逐字节一致 + 同一服务进程内。这对"固定素材 + 逐章/逐问"的场景非常对路:第 2 轮起 prefill 从 40 秒降到 2 秒以内,总耗时几乎等于纯 decode 时间。注意缓存只在进程存活期内有效,重启即丢。

一个容易踩的认知:中文 token 别用估算

我实测纯中文短文约 0.62~0.68 token/字符,但小说素材前缀 18k 字符却变成 23.9k token(≈1.33 token/字符),差一倍。文本类型差异太大,一律以服务端返回的 usage.prompt_tokens 为准。这跟本站的 Token 计数器 用途不同——那个工具适合粗略预算,精确计费必须看 API 返回值。

修复方案(已验证)

我写了一版优化脚本,改了三处,本地逻辑单测 21/21 通过,切到优化版后逐项复测:

修复项之前之后
max_tokens 类型校验掐断连接400 + 可读错误
max_tokens 越界(0/-1/超上限)泄露内部 C++ 错误400 + 范围提示 1~20480
空 messages 校验白生成 20s+ 废话400,不做生成
本地路径图片按来源限制任意文件读取仅 localhost 允许

改完之后 max_tokens="abc" 返回 400 max_tokens 格式错误: 'abc' 不是整数(合法范围 1~20480),空 messages 返回明确 400,远程调本地文件路径图片返回 400(只允许 base64),而正常文本、图片识别全部 200。并发问题按需求没改——单用户串行完全正常,要服务化再加前置队列就行。

如果你也要自建接口,检查清单

  • 必须有 /v1/models,否则一半客户端联不上。
  • 所有数值参数都要校验,非法值返回 400 而不是崩线程、不是泄露堆栈。
  • 管线加锁或排队,别假设底层会自己排队。
  • finish_reason 区分 stop 和 length,否则 tool-call 逻辑全错。
  • 不生效的参数要明确拒绝,静默忽略比报错更危险。
  • 文件路径类参数按来源限制,别留成任意文件读取。
数据说明以上全部为本机实测(235h / Arc 140T 核显 / OpenVINO GenAI / gemma-4-26B INT4),全部结论可用 Python 3.13 脚本复现。不同服务实现的坑点会有差异,本文列的是我这套的实际行为。

Everyone runs local models now, and wrapping one in an OpenAI-compatible endpoint (TuringSense, LiteLLM, various one-api projects) lets upstream apps connect with zero code changes. But "compatible" runs deep — I tested line by line and found 12 issues, 3 of which break production outright.

Subject: an OpenVINO GenAI service on my LAN box 235h, running gemma-4-26B (INT4, with a vision encoder). /health reports {"status":"ok","max_input_tokens":32000,"default_max_new":20480}. Every claim below is reproducible.

First, the good news: core paths all pass

ItemResult
Non-streaming chat/completionsPass, standard structure
Streaming (SSE)Pass, ends with [DONE]
Multi-turn messagesPass (remembers names)
system rolePass
max_tokens aliasmax_new_tokens also works
Image multimodalPass, quite accurate
Chinese UTF-8Pass throughout

Each streamed response carries a usage block (OpenAI only sends it with stream_options.include_usage) plus prompt_tps / decode_tps / prefill_ms / decode_ms. It's a superset — use it directly for monitoring instead of adding your own instrumentation.

Gotcha 1: concurrency causes instant 500 (worst)

Two simultaneous requests (12k tokens each): one returns 200, the other 500s with Generate cannot be called while ContinuousBatchingPipeline is already in running state.

Root cause: threaded HTTP + single pipeline instance

  • The HTTP layer is a ThreadingHTTPServer — it happily accepts concurrent requests.
  • But the underlying VLMPipeline.generate is a single global, non-thread-safe instance.
  • When two requests genuinely overlap, the second slams into a running pipeline → instant 500, not a queue.
What to doSerial calls are fine (send one, wait for it) — which is why it "looks single-threaded". But multiple clients hitting it at once will get random 500s and no retry means the request just dies. Fix: put a serializing queue or lock in front, or force serial + timeout retry on the client. Do not expect OpenVINO to queue for you.

Gotcha 2: bad max_tokens type kills the connection

Send "max_tokens": "abc" and the connection is simply severed — RemoteDisconnected, no error body. Clients just assume a network fault and retry.

Now try max_tokens=0 or -1: a 500 whose body leaks OpenVINO's internal C++ error and disk paths — C:\Jenkins\workspace\...pipeline_impl.cpp.

Cause: an uncaught ValueError from int(...) kills the handler thread. The fix is simple: validate input and unify error responses.

Gotchas 3-7: five compatibility gaps

3. Missing /v1/models breaks some clients

GET /v1/models returns 404. Many OpenAI clients and dashboards call client.models.list() during handshake, so they simply refuse to connect. The tested service only implements /v1/chat/completions and /health.

4. finish_reason is always length

Even "1+1=?" answered with "2" returns finish_reason: "length" (hardcoded). The docs call it intentional, but it's a trap: many SDKs branch on finish_reason == "stop" to decide completeness and tool-call handling, and they all misjudge.

5. temperature and friends are silently ignored

temperature / top_p / presence_penalty / seed / n / model are all silently ignored — no error, no effect, because the source never reads them. The client sends temperature=0.8 expecting sampling; the server is fixed greedy (do_sample=False). Silently ignoring is worse than erroring out, because users believe the setting took effect.

6. Images: only two accepted forms

An http(s):// image URL returns 400 (不支持 http(s) 图片链接) — fetching isn't implemented. Only base64 data URIs and server-local file paths work.

Security noteServer-local file paths mean arbitrary file read. I set the path to C:\Windows\web\wallpaper\Windows\img0.jpg and it read and described the Win11 wallpaper. Fine on a LAN, but a public deployment is a vulnerability. My fix: allow base64 from anywhere, and restrict local paths to localhost only.

7. Markdown image syntax fails badly

With markdown syntax ![img](data:image/png;base64,...), the server doesn't parse it and feeds the whole base64 blob in as plain text. In my test the model invented placeholder descriptions like "AABB CCDD…", burning 6k+ tokens. Some OpenAI client libraries pass images this way and will hit it.

Speed: in line with expectations

MetricMeasured
Cold prefill430~630 t/s
Decode12~16 t/s
20k input → first token~45 秒
20k prefix cached1.19 秒
Test summary for a self-hosted OpenAI-compatible API: 7 passes, 5 compatibility gaps, 2 serious bugs

Waiting 45 seconds for the first token on a cold 20k input is a crucial feel. But once the prefix cache hits, it's a completely different story:

CallTotalTTFTNote
1st (cold)45.1s44.5sFull prefill
2nd (same prefix, new question)1.19s0.98sCache hit
3rd (identical)0.58s0.44sFull cache

About 38x faster from cold to warm on a 20k prefix. It needs byte-identical prefixes in the same process. Perfect for "fixed material + per-chapter questions" workflows: from round 2, prefill drops from 40s to under 2s and total time ≈ pure decode. Cache dies with the process.

One more trap: don't estimate Chinese tokens

Plain Chinese prose measured 0.62–0.68 tokens/char, but a novel excerpt of 18k chars came out at 23.9k tokens (≈1.33/char) — a 2x difference. Text type matters hugely, so always trust the server's usage.prompt_tokens. This differs from our Token Counter, which is fine for rough budgeting but not for exact billing.

The fix (verified)

I wrote an optimized version touching three things. Local unit tests passed 21/21, and I re-verified each case on the new build:

FixBeforeAfter
max_tokens type checkConnection severed400 + 可读错误
Out of rangeLeaks internal C++ error400 + 范围提示 1~20480
Empty messages checkBurns 20s+ on nonsense400,不做生成
Local-path images by originArbitrary file read仅 localhost 允许

After the change, max_tokens="abc" returns 400 max_tokens 格式错误: 'abc' 不是整数(合法范围 1~20480), empty messages gets a clean 400, remote calls with local file paths are rejected (base64 only), and normal text and image recognition stay 200. Concurrency was left alone by design — serial use is fine; add a queue when you serve multiple users.

Checklist if you're building one

  • Ship /v1/models or half the clients won't connect.
  • Validate every numeric parameter — bad values must yield 400, not a dead thread or a leaked stack trace.
  • Lock or queue the pipeline — never assume the backend serializes for you.
  • Distinguish stop from length in finish_reason, or every tool-call path misfires.
  • Explicitly reject unsupported parameters — silent ignoring is more dangerous than an error.
  • Restrict file-path parameters by origin so you don't ship an arbitrary file read.
Data noteAll of this is measured on my own setup (235h / Arc 140T iGPU / OpenVINO GenAI / gemma-4-26B INT4) and reproducible with Python 3.13. Other server implementations will differ — these are the actual behaviors I observed.
返回文章列表