128G 统一内存 AI 工作站三选一:AMD 395 / DGX Spark / Mac Studio M5 Max

想在桌面上本地跑几百亿参数的大模型,三台 128G 统一内存的小机器是绕不开的选择:AMD 锐龙 AI Max+ 395、NVIDIA DGX Spark、Apple Mac Studio M5 Max。它们共用同一块 128GB 内存池,能装下 100B 级模型,但价格从 1.9 万到 4.5 万差出一倍多,跑同一个模型的输出速度更是从 13 到 27 tps 拉开档次。这篇把三台的定位、参数、实测速度、价格一次排开,数据全部来自官方口径与第三方评测汇总(2026.9 采集)。

一句话定位:AMD 395 = 性价比之王(Windows 全能);DGX Spark = 纯 AI 研究工具(CUDA 生态);Mac Studio M5 Max = 创意工作站(输出最顺)。

为什么这种小主机能跑大模型:统一内存

关键在统一内存。普通电脑内存归 CPU、显存归显卡,各管各的,跑大模型必须整个塞进显存,塞不下就跑不了;而统一内存是 CPU 和 GPU 共用一块内存池,128GB 全部共享。三台的「128GB」都是这个含义——内存即显存,模型想占多少占多少。

先记一个坑:统一内存让你装得下更大模型,但跑得快不快靠内存带宽而不是容量。容量回答「放不放得下」,带宽回答「放进去以后跑多快」——三台最核心的差距就在带宽。

核心参数总表(128GB 版)

项目AMD AI Max+ 395NVIDIA DGX SparkMac Studio M5 Max
CPU 架构Zen 5 (x86)Grace Blackwell (ARM)Apple M5 (ARM)
CPU16 核 32 线程20 核 (10×X925 + 10×A725)18 核 (6 超核 + 12 性能核)
GPURadeon 8060S (RDNA3.5, 40CU)Blackwell + 第5代 Tensor Core最高 40 核 (每核含神经加速)
AI 算力约 126 TOPS1000 TOPS FP4 / 500T FP8约 110 TOPS FP16
内存带宽约 256 GB/s约 273 GB/s614 GB/s
可分配显存最高 96GB统一 (全量)统一 (全量)
存储2TB+ 可扩展4TB 板载 NVMe1-8TB SSD
满载功耗约 120-140W约 240W约 200-270W
系统Windows / LinuxDGX OS (仅 Linux)macOS
参考价 (128GB)约 1.9-2.3 万约 3.5 万约 4.2-4.5 万

同一基准跑 Qwen3.5-27B:差距全在带宽

用同一个 Qwen3.5-27B(IQ4 量化)做统一基准,输出速度(decode)一测见真章。M5 Max 带宽是 AMD 395 的 2.4 倍、DGX Spark 的 2.25 倍,输出速度直接拉开一档:

Qwen3.5-27B 输出速度 (tps)
Mac Studio M5 Max~27
AMD AI Max+ 395~15
DGX Spark~13
内存带宽 (GB/s)
Mac Studio M5 Max614
DGX Spark273
AMD AI Max+ 395256

27 tps 接近正常阅读速度、基本流畅;13-15 tps 能出字但明显「结巴」,长时间读长内容会难受。注意 DGX Spark 有反杀:它 prefill(预填充/理解问题)特别快,提示吞吐能到 1957 tok/s(AMD 395 约 956),差不多是 Mac 的 3 倍——DGX 脑子转得快但说话结巴,Mac 说话最顺

价格速览(2026.9 京东实拍)

产品配置到手价
AMD 阿迈奇 MIA PRO+128G+0T¥19,402
AMD E人E本 G95128G 无盘¥19,899
AMD TOPCAMD128G¥19,889
AMD 极摩客 EVO-X2128G+2T¥22,874
华硕 GX10 (DGX OEM)128GB+1TB¥34,824
NVIDIA DGX Spark (FE)128GB+4TB¥38,999
Mac Studio M5 Max128GB+1TB¥41,749
Mac Studio M5 Max128GB+2TB约 ¥45,499

受内存涨价影响,这类机器 2026 年初到 9 月普遍上涨:AMD 128GB 版从 1.3-1.5 万涨到 1.9-2.3 万,DGX Spark 年后官方涨了 17.5%(+700 美元),Apple 则下架了 M3 Ultra 的大内存选项。现在买 AI 主机确实不如年初划算,不急可以再等等。

各自短板(没完美的机器)

  • AMD 395:带宽 256 GB/s 输出偏慢(27B 约 15 tps);ROCm 生态不如 CUDA;板载内存不可扩展,64GB 版是「重灾区」。
  • DGX Spark:只支持 Linux(DGX OS),跑不了 3A 游戏;带宽 273 GB/s 输出偏慢(27B 约 13 tps);涨价后逼近 4 万性价比受质疑;多机集群配置繁琐。
  • Mac Studio M5 Max:贵(比 AMD 395 贵约 1.8 万);不支持 CUDA/vLLM,批处理与微调弱(LoRA 比 DGX 慢约 4 倍,70B 全量微调不可行);MLX 生态不如 CUDA 成熟。

到底怎么选

  • 预算有限、想本地跑大模型、兼顾办公/游戏 → AMD AI Max+ 395(128GB 最便宜、Windows 全能,输出慢点能忍)
  • 专业 AI 研发、微调、批处理、离不开 CUDA → NVIDIA DGX Spark(算力最强、生态最全、可扩展,但只支持 Linux)
  • 创意工作者、追求流畅本地推理、macOS 生态 → Mac Studio M5 Max(带宽最高、输出最顺、全能工作站,就是贵)
  • 偶尔玩玩、不是重度生产 → 其实不急着上这几万级的,先看独显整机或等降价更实在
只给一个人用 → M5 Max;团队/多 Agent 并发 → 双 DGX Spark;预算紧张又要 Windows 全能 → AMD 395。
重要提醒:统一内存不可扩展,买错容量只能整机换。这几台务必一步到位买 128GB 版,64GB 是重灾区。

数据说明:本文速度/跑分数据来自 AMD 官方 ROCm 数据、NVIDIA 官方宣传与知乎/什么值得买/IT之家/Notebookcheck/至顶网等第三方评测汇总(2026.9 采集),不同测试环境有差异,仅供参考;价格为 2026-09-01 京东实时到手价,会随市场波动。

If you want to run 100B-scale LLMs locally on a desktop, three 128GB unified-memory machines keep coming up: AMD Ryzen AI Max+ 395, NVIDIA DGX Spark, and Apple Mac Studio M5 Max. They share one pool of 128GB memory, so they can fit 100B-class models — yet prices span ¥19k-45k and decode speed on the same model ranges from 13 to 27 tps. This post lays out positioning, specs, measured speed and prices for all three, sourced from official claims and third-party reviews (collected Sept 2026).

One-liner: AMD 395 = best value (all-round Windows box); DGX Spark = pure AI research tool (CUDA ecosystem); Mac Studio M5 Max = creative workstation (smoothest output).

Why these tiny boxes run big models: unified memory

The key is unified memory. On a normal PC, RAM belongs to the CPU and VRAM to the GPU; an LLM must fit entirely in VRAM or it won't run. With unified memory, CPU and GPU share one pool — all 128GB is usable by the GPU. On all three machines '128GB' means exactly that: memory is VRAM.

Remember:Unified memory lets you fit bigger models, but speed is decided by memory bandwidth, not capacity. Capacity asks 'does it fit?'; bandwidth answers 'how fast does it run once loaded'. Bandwidth is exactly where these three differ.

Specs at a glance (128GB)

SpecAMD AI Max+ 395NVIDIA DGX SparkMac Studio M5 Max
CPU archZen 5 (x86)Grace Blackwell (ARM)Apple M5 (ARM)
CPU16C/32T20C (10×X925 + 10×A725)18C (6 perf + 12 eff)
GPURadeon 8060S (RDNA3.5, 40CU)Blackwell + 5th-gen Tensor Coresup to 40-core + neural accel
AI compute~126 TOPS1000 TOPS FP4 / 500T FP8~110 TOPS FP16
Bandwidth~256 GB/s~273 GB/s614 GB/s
GPU-usable RAMup to 96GBunified (all)unified (all)
Storage2TB+ expandable4TB onboard NVMe1-8TB SSD
Max power~120-140W~240W~200-270W
OSWindows / LinuxDGX OS (Linux only)macOS
Price (128GB)~¥19k-23k~¥35k~¥42k-45k

Same Qwen3.5-27B benchmark: it's all bandwidth

Under one uniform benchmark — Qwen3.5-27B (IQ4 quantized) — decode speed tells the whole story. The M5 Max has 2.4× the bandwidth of the AMD 395 and 2.25× the DGX Spark, and decode speed opens up a full tier:

Qwen3.5-27B decode (tps)
Mac Studio M5 Max~27
AMD AI Max+ 395~15
DGX Spark~13
Memory bandwidth (GB/s)
Mac Studio M5 Max614
DGX Spark273
AMD AI Max+ 395256

27 tps is close to reading speed and feels smooth; 13-15 tps produces text but stutters, which gets tiring for long outputs. DGX Spark fights back on prefill (understanding your prompt): prompt throughput reaches 1957 tok/s (AMD 395 ~956), roughly 3× the Mac — DGX thinks fast but talks with a stutter; the Mac talks smoothest.

Street prices (JD, Sept 2026)

ProductConfigStreet price
AMD Amaic MIA PRO+128G no disk¥19,402
AMD E人E本 G95128G no disk¥19,899
AMD TOPCAMD128G¥19,889
AMD GMKtec EVO-X2128G+2T¥22,874
Asus GX10 (DGX OEM)128GB+1TB¥34,824
NVIDIA DGX Spark (FE)128GB+4TB¥38,999
Mac Studio M5 Max128GB+1TB¥41,749
Mac Studio M5 Max128GB+2TB~¥45,499

DRAM price hikes pushed all three up between early and late 2026: AMD 128GB units went from ¥13-15k to ¥19-23k, NVIDIA raised DGX Spark 17.5% (+$700), and Apple dropped the big-RAM M3 Ultra options. Buying AI hardware now is pricier than in January — if you're not in a hurry, waiting pays.

The catch with each

  • AMD 395: 256 GB/s bandwidth = slow decode (~15 tps on 27B); ROCm ecosystem lags CUDA; memory is soldered, and the 64GB version is the one to avoid.
  • DGX Spark: Linux (DGX OS) only, no 3A gaming; 273 GB/s = slow decode (~13 tps on 27B); post-hike price near ¥40k hurts value; multi-node clustering is fiddly.
  • Mac Studio M5 Max: expensive (~¥18k more than the AMD); no CUDA/vLLM, weak batch + finetuning (LoRA ~4× slower than DGX, no full 70B tuning); MLX ecosystem less mature than CUDA.

How to choose

  • On a budget, want local LLMs plus office/gaming → AMD AI Max+ 395 (cheapest 128GB, all-round Windows, slower output is tolerable)
  • Professional AI research, finetuning, batch jobs, need CUDA → NVIDIA DGX Spark (most compute, full ecosystem, expandable, Linux only)
  • Creative pro, want smooth local inference, in the macOS world → Mac Studio M5 Max (highest bandwidth, smoothest output, all-round workstation, just pricey)
  • Occasional tinkering, not heavy production → you don't need a ¥20-45k box; a discrete-GPU build or waiting for price drops makes more sense
One user → M5 Max; team / many agents → dual DGX Spark; tight budget but want an all-round Windows box → AMD 395.
Heads-up:Unified memory can't be upgraded — a wrong capacity means a whole new machine. Go straight for the 128GB models; 64GB is the version to avoid.

Data notes: speed figures come from AMD official ROCm numbers, NVIDIA marketing and third-party reviews (Zhihu, SMZDM, ITHome, Notebookcheck, ZDNet, etc., collected Sept 2026); results vary by test env and are for reference only. Prices are JD.com street prices as of 2026-09-01 and fluctuate.

返回文章列表