Mac mini 能跑多大的本地模型? How Big a Local Model Can Your Mac mini Run?

用 Qwen 3.5 六档模型的真实下载体积,讲清统一内存的容量边界、适用任务和验收方法。数据于 2026-10-07 从 Ollama 官方模型库实时核实。 Using the real download sizes of six Qwen 3.5 tiers, this explains the memory ceiling, what each tier is actually good for, and how to benchmark it yourself. Figures verified live from the Ollama library on 2026-10-07.

先给结论The Short Answer

「模型文件能下载,机器就能跑」——这是本地模型最容易误导人的一句话。 "If the file downloads, the machine can run it" — the single most misleading claim in local AI.

模型权重只是第一笔内存开销。系统、推理引擎、上下文缓存、图像编码器和并发请求,全都要从同一块统一内存里分走一杯羹。所以容量问题从来不是「最大能启动多少 B」,而是「在你要用的上下文长度和任务下,能稳定跑到多少 B」。 Model weights are only the first claim on memory. The OS, the inference engine, the KV cache, the vision encoder and concurrent requests all draw from the same unified pool. So the real question is never "what's the biggest model that boots" — it's "what runs stably at the context length and workload you actually need."

一、模型体积 ≠ 运行内存1. File Size Is Not Memory Footprint

Ollama 列出的下载体积,描述的是量化后的权重文件大小,不是运行时的内存占用上限。真正吃内存的还有四样东西: The download size listed by Ollama describes the quantized weight file — not the runtime memory ceiling. Four more things compete for the same pool:

上下文缓存KV cache

对话越长,缓存越大。这是最容易被忽略、也最容易把内存撑爆的一项。The longer the conversation, the larger the cache. Easiest to overlook, easiest to blow your budget on.

视觉能力Vision encoder

加载图像输入会额外占用内存,Qwen 3.5 本身是多模态模型,图像输入会显著抬高占用。Image input costs extra memory. Qwen 3.5 is multimodal, so vision pushes the footprint up noticeably.

工具调用与 AgentTools & agents

搜索、代码工具类任务往往希望更长上下文,会进一步压缩内存余量。Search and coding agents want longer context, which squeezes the remaining headroom further.

并发会话Concurrent sessions

多个会话同时跑,内存是叠加的。单人使用和多人共用完全是两个容量问题。Sessions stack. Single-user and shared-server are two entirely different capacity problems.

二、六档模型实测数据2. The Six Tiers, Verified

下表「下载体积」为 2026-10-07 从 Ollama 官方模型库核实的默认量化版本,非文章原述照搬。 "Download" below is the default quantized variant verified from the Ollama library on 2026-10-07 — not copied from any article.

档位Tier 运行命令Command 下载体积Download 定位Role 32GB 适用性On 32GB
0.8B qwen3.5:0.8b 1.2 GB 自动化零件Automation part 轻量常驻Always-on
2B qwen3.5:2b 1.9 GB 极轻量日常Ultra-light daily 余量充足Plenty spare
4B qwen3.5:4b 3.3 GB 日常轻任务起点Entry daily driver 舒适Comfortable
9B qwen3.5:9b 6.6 GB 最值得长期使用Best long-term pick 最均衡Sweet spot
27B qwen3.5:27b 17 GB 能力优先单任务One task, max quality 任务档,非常驻Task mode only
35B qwen3.5:35b 22 GB 接近容量边界At the ceiling 可尝试,勿作主力Try, don't rely
122B qwen3.5:122b 81 GB 远超单机范围Far beyond desktop 不可行Not viable

三、每档到底拿来干嘛3. What Each Tier Is Actually For

0.8B — 自动化零件,不是聊天助手0.8B — a pipeline part, not a chatbot

ollama run qwen3.5:0.8b

优势是轻、快、容易常驻,不是回答复杂问题。适合:文本打标签、提取固定字段、判断是否需要人工处理、生成简单文件名、自动化流程第一轮分流。Wins on being light, fast and always resident — not on hard questions. Good for: tagging text, extracting fixed fields, triage decisions, generating filenames, first-pass routing in a pipeline.

⚠️ 小模型被要求做复杂分析时,常见结果不是明确拒绝,而是很自信地编。用在流水线里可以,当全能助手不行。⚠️ Ask a small model for complex analysis and it rarely refuses — it confidently makes things up. Fine inside a pipeline; not as a general assistant.

4B — 日常轻任务的起点4B — where daily light work starts

ollama run qwen3.5:4b

适合摘要、改写、会议记录整理和简单图片理解。留下的内存余量充足,更容易同时跑 Web UI、浏览器和其他服务。Summary, rewrite, meeting-note cleanup, simple image understanding. Leaves enough headroom to also run a web UI, a browser and other services.

如果目标是让机器长期在线、提供轻量 AI 入口,这一档往往比「大到刚好塞满」更稳定。If the goal is a box that stays up and serves light AI, this tier is usually more reliable than filling memory to the brim.

9B — 最值得长期使用的一档 ★9B — the best long-term pick ★

ollama run qwen3.5:9b

能力比小模型完整,又不会把统一内存逼到边缘。适合:较长文档摘要、中文写作与结构调整、本地知识库问答、常规代码解释、单人使用的工具调用。Meaningfully more capable than the small tiers without pushing memory to the edge. Good for: longer document summaries, writing and restructuring, local knowledge-base Q&A, everyday code explanation, single-user tool calls.

如果不知道从哪一档开始:先用 9B 完成真实任务,再决定是否上 27B。If unsure where to start: run 9B on your real work first, then decide whether 27B is worth it.

27B — 能跑,但要开始做取舍27B — runs, but now you're trading off

ollama run qwen3.5:27b

默认量化版本 17GB。32GB 统一内存可以容纳权重并为系统和上下文留下空间——但这不等于可以随意开超长上下文、多个并发对话,再同时跑一堆桌面软件。The default quant is 17GB. On 32GB unified memory the weights fit with room for the OS and context — but that does not mean you can run huge context, several conversations and a pile of desktop apps at once.

这一档适合单用户、单模型、较复杂的写作和推理。使用建议:Suited to single-user, single-model, heavier writing and reasoning. Practical rules:

  • 先关闭不必要的大型应用Close unnecessary heavy apps first
  • 从较短上下文开始Start with a shorter context
  • 不同时常驻多个大模型Don't keep several large models resident at once
  • 用 ollama ps 检查是否完整使用 GPUUse ollama ps to confirm full GPU offload
  • 看内存压力,而不是只看剩余内存数字Watch memory pressure, not just the free-memory number

27B 更像「任务档」,不是最舒服的「全天常驻档」。27B is a task-mode tier, not a comfortable always-on tier.

35B — 文件放得下,不等于值得跑35B — it fits, but that doesn't make it worth running

ollama run qwen3.5:35b

默认版本 22GB,权重已吃掉大部分统一内存。它可能在短上下文、单任务条件下启动,但这不代表能获得稳定顺滑的长期体验。只要上下文增加、图像输入变大,或后台程序争抢内存,就会出现明显等待和内存压力。The default build is 22GB — most of the pool gone before the OS and cache claim theirs. It may boot under short context and a single task, but that is not a stable long-term experience. Add context, larger images, or background apps competing for memory, and you'll see real stalls.

如果长期主力需求就是 35B,应该上更大统一内存,而不是把现有机器当极限去挑战。If 35B is your long-term daily need, buy more unified memory — don't make this machine the thing you fight against.

四、16GB 机器怎么办(你的实际情况)4. What About a 16GB Machine? (Your Case)

以上按 32GB 讲。实测本机为 16GB 统一内存 / Apple M4,容量结论必须下修——这不是打折,而是换个甜点档。 The above assumes 32GB. This machine is 16GB unified / Apple M4, so the conclusions shift down — not a discount, just a different sweet spot.

档位Tier 体积Size 16GB 判定Verdict on 16GB 说明Notes
0.8B / 2B 1.2–1.9 GB ✓ 轻松Easy 随便跑,可长期后台常驻Run freely, fine to keep resident
4B 3.3 GB ✓ 舒适Comfortable 甜点档,可同时跑浏览器等常用软件The sweet spot; browser and normal apps coexist
9B 6.6 GB ~ 可行Workable 能跑,但长上下文要收敛;跑前关大应用Runs, but keep context modest; close heavy apps first
27B 17 GB ✗ 装不下Won't fit 权重已超过系统占用后剩下的可用内存Weights exceed memory left after the OS
35B 22 GB ✗ 不可行No 远超容量,不要尝试Far past capacity; don't bother

⚠️ 为什么 16GB 的上限是 9B 而不是 27B:macOS 自身与常驻服务通常会占掉数 GB,再扣掉上下文缓存,留给模型权重的余量大致在 11–13GB 区间。27B 的 17GB 权重单这一项就超了。这个区间随你后台开了什么而浮动,属于工程估算,无法用单一数字精确交叉验证。 ⚠️ Why 16GB tops out at 9B, not 27B: macOS and resident services typically consume several GB, and the KV cache takes more, leaving roughly 11–13GB for weights. The 27B build's 17GB of weights alone blows that. The exact range moves with whatever is running — this is an engineering estimate, not a single verifiable figure.

五、别抄别人的速度,自己跑五项验收5. Don't Copy Benchmarks — Run Your Own Five Checks

同一台机器上,真正有意义的只有横向对比。所以不要迷信任何文章里的「每秒多少 token」——不同芯片、量化版本、上下文和温度都会改变结果。每档模型都跑下面五项: On your own machine, only a side-by-side comparison means anything. So don't trust any published "tokens per second" — chip, quantization, context and temperature all move the number. Run all five checks per tier:

  1. 首次回答要等多久How long does the first answer take?

  2. 连续生成时是否明显卡顿Does continuous generation visibly stutter?

  3. 任务事实是否正确Are the facts in the answer correct?

  4. 长对话后是否开始遗忘或混乱Does it start forgetting or confusing things in long chats?

  5. 运行时系统还能否正常使用Can you still use the machine normally while it runs?

配套命令The companion commands

ollama ps

查看运行中的模型与 GPU 占用情况。Shows what's loaded and whether it's on the GPU.

open -a "Activity Monitor"

看内存压力(Memory Pressure),不要只看剩余数字。Watch the Memory Pressure graph — not just the free-memory figure.

ollama stop qwen3.5:27b

测完立即卸载当前模型,再加载下一档,避免多模型同时占内存污染比较结果。Unload immediately after testing, then load the next tier — otherwise two models share memory and the comparison is worthless.

六、用什么任务测试最公平6. What to Test With

不要只问数学题,也不要只看它能不能写一首诗。准备一套与你真实工作有关的固定题目: Don't just ask math questions, and don't judge it on whether it can write a poem. Prepare a fixed set drawn from your actual work:

一份 3000 字会议记录,提取行动项A 3,000-word meeting transcript → extract action items

两份产品说明,比较差异Two product specs → compare the differences

一段包含错误的代码,解释并修复A snippet with bugs → explain and fix

一张带表格的截图,提取字段A screenshot with a table → extract the fields

一个资料中没有答案的问题,测是否会乱编A question with no answer in the source → does it hallucinate?

把回答时间、正确项、错误项和内存压力记在同一张表里。模型选择不是比谁参数大,而是比谁能在你的机器上稳定交付。 Log response time, hits, misses and memory pressure in one table. Choosing a model isn't about parameter count — it's about what delivers reliably on your hardware.

七、数据核实与更正7. Verification & Corrections

本页所有体积数据于 2026-10-07 从 Ollama 官方模型库(ollama.com/library/qwen3.5)实时抓取核实,与流传版本的原文存在以下出入: All size figures on this page were pulled live from the official Ollama library (ollama.com/library/qwen3.5) on 2026-10-07. The following discrepancies exist against the widely circulated original:

✓

核对无误:9B = 6.6GB、27B = 17GB、4B = 3.3GB(默认量化版本,与原文一致)Confirmed accurate: 9B = 6.6GB, 27B = 17GB, 4B = 3.3GB (default quants, matching the original)

⚠

数据偏差:0.8B 原文称约 1.0GB,实测默认版本为 1.2GB;1.0GB 对应的是 q8_0 或 nvfp4 量化变体,不是默认拉取版本。Discrepancy: the original cites ~1.0GB for 0.8B; the actual default is 1.2GB. The 1.0GB figure belongs to the q8_0 / nvfp4 variants, not what ollama run pulls by default.

⚠

数据偏差:35B 原文称约 24GB,实测默认(a3b q4_K_M)为 22GB;24GB 对应 mtp-q4_K_M 变体。Discrepancy: the original cites ~24GB for 35B; the actual default (a3b q4_K_M) is 22GB. 24GB corresponds to the mtp-q4_K_M variant.

+

原文遗漏:Qwen 3.5 家族还有 2B(1.9GB)与 122B(81GB)两档未被提及。2B 是比 0.8B 更实用的轻量档,122B 则完全超出单机范围。Omitted by the original: the Qwen 3.5 family also includes 2B (1.9GB) and 122B (81GB). 2B is a more useful light tier than 0.8B; 122B is far beyond any single desktop.

+

模型性质补充:Qwen 3.5 是多模态开源模型家族(支持文本与图像输入),上下文窗口 256K。原文提到「视觉能力占内存」是对的,但没说清这是模型原生能力。Model nature: Qwen 3.5 is a multimodal open-source family (text + image input) with a 256K context window. The original's point about vision costing memory is correct, but it never states that vision is native to the model.

?

无法交叉验证:各档在具体芯片上的实际生成速度(tokens/s)。原文明确声明不编造这组数字,这一处理是诚实的——速度受芯片、量化、上下文、温度影响,任何单一数字都不可迁移。Not verifiable: actual generation speed (tokens/s) per tier on specific chips. The original explicitly declines to fabricate these figures — an honest call, since speed varies with chip, quantization, context and temperature, and no single number transfers.

买机器前该问自己的四个问题Four Questions Before You Buy

最该问的不是「最大能启动多少 B」,而是: The question isn't "what's the biggest model that boots." It's:

① 我每天真正做什么任务? ② 我需要多长上下文?
③ 我是单人使用,还是多人并发? ④ 我愿不愿意为了更大模型牺牲速度和系统余量?
① What do I actually do every day? ② How much context do I need?
③ Single user or concurrent? ④ Will I trade speed and system headroom for a bigger model?

如果答案只是摘要、资料问答和常规写作,9B 级已经很实用。如果经常处理复杂长文档和代码,27B 值得考虑——但它需要比 32GB 更宽裕的内存才谈得上长期舒适。如果长期盯着 35B 以上,就不要让一台勉强能启动它的机器,去承担一件它很难舒服完成的工作。 If the answers are summary, document Q&A and routine writing, 9B is already practical. If heavy long documents and code are routine, 27B is worth considering — but comfortable long-term use wants more memory than 32GB. And if 35B-plus is your permanent target, don't ask a machine that barely boots it to do work it can't comfortably finish.