1. 为什么 MCP Sampling 值得你花时间配置
MCP Sampling 是 Model Context Protocol 里最容易被忽略、但实际价值极高的一块能力。简单说,它让 MCP Server 可以主动向 Client 发起一次 LLM 采样请求,而 Client 在真正调用模型之前,可以调整参数、甚至人工修改结果。这跟传统的「Server 定义提示词、Client 直接执行」完全不同——主动权在 Server,但控制权在 Client。
适合谁?如果你正在用 Claude Code、Cline、Codex 这类本地 AI 工具,并且希望在某些关键生成环节插入人工确认,或者想针对不同任务动态调整 temperature、maxTokens,那 Sampling 就是你要找的东西。它解决的核心痛点是:LLM 生成过程不可控、不可审计、不可干预。
我试过把 Sampling 用在文件系统助手上,Server 端发起采样请求,Client 端弹出参数调整和结果确认,整个链路跑通后,生成质量的可控性提升非常明显。下面我会以 settings.json 为骨架,把参数微调和人工干预开关的协同配置完整拆一遍,所有配置片段都可以直接复制。
在开始之前,你需要一个统一的 API 通道来管理 Key 和模型调用。TaoToken 提供了统一的 Key/API 通道,官网入口是 https://taotoken.net/?utm_source=taotoken_aicg_blog_end&utm_medium=csdn&utm_campaign=rewrite&utm_content= ,API 地址是 https://taotoken.net/api 。后面所有配置里的 Base URL 都指向这里,你只需要在控制台生成一个 Key 即可。
2. TaoToken 前置准备:Key、Base URL 与模型 ID
在写 settings.json 之前,先把三件套准备好:Base URL、API Key、Model ID。这三样东西贯穿整个 Sampling 配置,缺一个都跑不起来。
Base URL 固定为https://taotoken.net/api,注意这里不加任何 UTM 参数,保持干净。API Key 需要你去控制台生成,入口在 https://taotoken.net/console?utm_source=taotoken_aicg_blog_end&utm_medium=csdn&utm_campaign=rewrite&utm_content= 。生成后复制出来,后面填到 settings.json 的 env 字段里。
Model ID 这块要特别注意。Sampling 请求里的modelPreferences.hints填的是模型名称,但 Client 端真正调用时用的 Model ID 必须和 TaoToken 支持的模型列表对齐。常见的比如claude-sonnet-4-20250514、gpt-4o、deepseek-chat这些都可以。你可以在模型对话页面先验证一下模型是否可用,入口是 https://taotoken.net/chat?utm_source=taotoken_aicg_blog_end&utm_medium=csdn&utm_campaign=rewrite&utm_content= 。
如果你打算长期做编码类 Agent 开发,建议直接开 Coding Plan,入口在 https://taotoken.net/coding-plan?utm_source=taotoken_aicg_blog_end&utm_medium=csdn&utm_campaign=rewrite&utm_content= ,这样采样请求的调用配额会更充裕。
API Key 的管理页面在 https://taotoken.net/api-keys?utm_source=taotoken_aicg_blog_end&utm_medium=csdn&utm_campaign=rewrite&utm_content= ,接入文档在 https://taotoken.net/doc?utm_source=taotoken_aicg_blog_end&utm_medium=csdn&utm_campaign=rewrite&utm_content= 。Claude Code 相关的接入说明在 https://taotoken.net/claudecode-anthropic?utm_source=taotoken_aicg_blog_end&utm_medium=csdn&utm_campaign=rewrite&utm_content= 。
把这三样东西记下来,下面直接进配置。
3. 可复制配置:settings.json 完整片段
这一节是全文的核心。我会给出一个完整的 settings.json 片段,包含 MCP Server 定义、Sampling 参数、人工干预开关三部分。你可以直接复制到你的项目里,改掉 Key 和路径就能用。
先看整体结构。settings.json 的顶层是mcpServers,每个 Server 一个条目。Sampling 相关的配置放在env和sampling两个字段里。env负责注入 Base URL 和 Key,sampling负责控制采样行为和人工干预。
{ "mcpServers": { "file-system-assistant": { "command": "python", "args": ["server/server.py"], "cwd": "./mcp-sampling-demo", "env": { "TAOTOKEN_BASE_URL": "https://taotoken.net/api", "TAOTOKEN_API_KEY": "sk-your-key-here", "TAOTOKEN_MODEL_ID": "claude-sonnet-4-20250514" }, "sampling": { "enabled": true, "defaultTemperature": 0.7, "defaultMaxTokens": 1000, "modelPreferences": { "hints": [ { "name": "claude-sonnet-4-20250514" }, { "name": "gpt-4o" } ], "costPriority": 0.5, "speedPriority": 0.7, "intelligencePriority": 0.8 }, "humanIntervention": { "beforeSampling": true, "afterSampling": true, "allowParamAdjust": true, "allowResultEdit": true, "timeoutSeconds": 120 }, "stopSequences": ["\n\n", "---"], "includeContext": "thisServer" } } } }逐字段说明。env.TAOTOKEN_BASE_URL固定填https://taotoken.net/api,这是所有 LLM 调用的统一入口。env.TAOTOKEN_API_KEY填你在控制台生成的 Key。env.TAOTOKEN_MODEL_ID是默认模型,Client 端在没有收到 hints 时会用它。
sampling.enabled是总开关,设为 true 才会处理采样请求。defaultTemperature和defaultMaxTokens是兜底值,当 Server 端没有指定时使用。modelPreferences里的三个 priority 参数控制模型选择策略,值域 0 到 1,越大表示越看重该维度。
humanIntervention是人工干预的核心配置。beforeSampling设为 true 时,Client 会在调用 LLM 之前暂停,把采样请求展示给用户,允许调整参数。afterSampling设为 true 时,LLM 返回结果后会再次暂停,允许用户修改结果。allowParamAdjust和allowResultEdit分别控制这两个阶段是否允许编辑。timeoutSeconds是等待用户输入的超时时间,超时后走默认值。
stopSequences是停止序列,遇到这些字符串就停止生成。includeContext设为thisServer表示把当前 Server 的上下文包含进去。
如果你用的是 TOML 格式的配置(比如某些 Codex 场景),等价写法如下:
[mcp_servers.file-system-assistant] command = "python" args = ["server/server.py"] cwd = "./mcp-sampling-demo" [mcp_servers.file-system-assistant.env] TAOTOKEN_BASE_URL = "https://taotoken.net/api" TAOTOKEN_API_KEY = "sk-your-key-here" TAOTOKEN_MODEL_ID = "claude-sonnet-4-20250514" [mcp_servers.file-system-assistant.sampling] enabled = true defaultTemperature = 0.7 defaultMaxTokens = 1000 stopSequences = ["\n\n", "---"] includeContext = "thisServer" [mcp_servers.file-system-assistant.sampling.humanIntervention] beforeSampling = true afterSampling = true allowParamAdjust = true allowResultEdit = true timeoutSeconds = 120配置写完后,Server 端发起采样请求时,Client 会读取这些字段。Server 端的采样请求结构长这样:
sampling_request = { "method": "sampling/createMessage", "params": { "messages": [ { "role": "user", "content": {"type": "text", "text": question} } ], "modelPreferences": { "hints": [{"name": "claude-sonnet-4-20250514"}], "costPriority": 0.5, "speedPriority": 0.7, "intelligencePriority": 0.8 }, "systemPrompt": "你是一个专业的文件系统助手。", "temperature": 0.7, "maxTokens": 1000, "stopSequences": ["\n\n"], "includeContext": "thisServer", "metadata": {"requestType": "file-system-query"} } }注意 Server 端的temperature和maxTokens会覆盖 settings.json 里的默认值,但 Client 端在人工干预阶段可以再次调整。这就是「参数微调」和「人工干预」的协同点:Server 给建议值,Client 做最终决策。
4. 验证请求:一次完整的采样链路
配置写好后,必须验证整条链路能跑通。这一节给出可执行的验证步骤和预期结果。
先启动 Server。在终端里执行:
cd mcp-sampling-demo python server/server.pyServer 启动后会打印「文件系统助手已启动,等待连接...」。然后另开一个终端启动 Client:
cd mcp-sampling-demo python client/client.py ../server/server.pyClient 连接成功后会列出可用的提示模板。输入一个问题,比如「请解释什么是 inode」,Client 会收到采样请求并展示出来:
============================================================ 服务器发送的采样请求: ============================================================ 方法: sampling/createMessage 系统提示: 你是一个专业的文件系统助手。 温度: 0.7 (0=保守, 1=创意) 最大令牌数: 1000 推荐模型: [{'name': 'claude-sonnet-4-20250514'}] 是否要修改采样参数?(y/n) > y 请输入新的 temperature (0.0-1.0,默认0.7): > 0.3 请输入新的 maxTokens (默认1000): > 500参数调整后,Client 调用 TaoToken 的 API 执行采样。调用时用的 Base URL 是https://taotoken.net/api,Key 从环境变量读取,Model ID 从 hints 里取。请求体大致如下:
response = client.chat.completions.create( model="claude-sonnet-4-20250514", messages=[ {"role": "system", "content": "你是一个专业的文件系统助手。"}, {"role": "user", "content": "请解释什么是 inode"} ], temperature=0.3, max_tokens=500 )LLM 返回结果后,Client 展示给用户:
============================================================ LLM采样完成,显示结果... ============================================================ 采样结果: { "model": "claude-sonnet-4-20250514", "stopReason": "endTurn", "role": "assistant", "content": { "type": "text", "text": "inode 是文件系统中用于存储文件元数据的数据结构..." } } 是否要修改结果?(y/n) > n如果选择修改,可以手动编辑文本,修改后的结果会作为最终采样结果返回给 Server。整个链路验证通过后,你会看到 Server 端收到最终结果并继续后续逻辑。
验证要点有三个:第一,采样请求的 JSON 格式必须符合 MCP 协议,method字段是sampling/createMessage;第二,参数调整后 LLM 确实使用了新参数,可以通过对比 temperature 0.3 和 0.7 的输出差异来确认;第三,人工修改的结果确实被返回给 Server,而不是被丢弃。
如果验证过程中出现401错误,说明 API Key 有问题,去 https://taotoken.net/api-keys?utm_source=taotoken_aicg_blog_end&utm_medium=csdn&utm_campaign=rewrite&utm_content= 重新生成一个。如果出现local proxy failed,检查 Base URL 是否写成了https://taotoken.net/api,不要多加斜杠或路径。
5. 常见错误排查:401、local proxy failed、reading choices、OAuth
这一节对照真实报错,给出排查路径。每个错误都给出触发场景、根因和修复方法。
401 Unauthorized。触发场景:Client 调用 LLM 时返回 401。根因通常是 API Key 无效、过期或没填。检查 settings.json 里的TAOTOKEN_API_KEY是否以sk-开头,是否有多余空格。如果 Key 是从控制台复制的,确认没有复制到换行符。修复方法:去 https://taotoken.net/api-keys?utm_source=taotoken_aicg_blog_end&utm_medium=csdn&utm_campaign=rewrite&utm_content= 重新生成,替换后重启 Client。
local proxy failed。触发场景:Client 启动时报连接失败。根因是 Base URL 配置错误,或者网络层无法到达https://taotoken.net/api。检查 settings.json 里的TAOTOKEN_BASE_URL是否精确等于https://taotoken.net/api,不要写成https://taotoken.net/api/v1或带尾部斜杠。修复方法:改成标准地址,重启。
reading choices 报错。触发场景:LLM 返回结果解析失败,报reading 'choices'或类似字段缺失。根因是 API 返回结构不符合预期,通常是因为 Model ID 写错了,或者请求体里多传了模型不支持的参数(比如某些模型不支持stopSequences)。修复方法:检查TAOTOKEN_MODEL_ID是否在 TaoToken 支持列表里,去掉不支持的参数。可以在模型对话页面先验证模型可用性:https://taotoken.net/chat?utm_source=taotoken_aicg_blog_end&utm_medium=csdn&utm_campaign=rewrite&utm_content= 。
OAuth 相关报错。触发场景:某些工具在接入时走 OAuth 流程失败。根因是认证方式不匹配。TaoToken 的接入用的是 API Key 方式,不需要 OAuth。如果你在 Claude Code 或 Codex 里看到 OAuth 报错,检查是否误开了 OAuth 模式。修复方法:在配置里显式指定 API Key 认证,参考接入文档 https://taotoken.net/doc?utm_source=taotoken_aicg_blog_end&utm_medium=csdn&utm_campaign=rewrite&utm_content= 。
另外,如果你用的是 CC Switch 或 Cline MCP,配置里必须写全三件套:Base URL、Key、Model ID。缺任何一个都会导致采样请求失败。CC Switch 的配置路径通常在~/.cc-switch/config.json,Cline MCP 的配置在 VS Code 的 settings.json 里。Codex 的 auth.json 里需要填api_key和base_url两个字段。
排查时建议打开 Client 的调试日志,把采样请求和 LLM 响应都打印出来。大部分问题看一眼原始 JSON 就能定位。
6. 把 Sampling 用起来:从验证到日常
配置跑通之后,Sampling 的真正价值在于日常使用中的参数微调和人工干预。你可以针对不同任务预设不同的采样策略。比如代码生成用 temperature 0.2、maxTokens 2000,创意写作用 temperature 0.8、maxTokens 1500,问答系统用 temperature 0.5、maxTokens 500。
人工干预开关也不是一直开着就好。批量任务时可以关掉beforeSampling和afterSampling,让流程自动跑;关键内容生成时再打开,确保每一步都可控。timeoutSeconds建议设成 120 秒,太短容易误触默认值,太长会卡住流程。
如果你要做更复杂的 Agent 开发,可以把 Sampling 和 Coding Plan 结合,入口在 https://taotoken.net/coding-plan?utm_source=taotoken_aicg_blog_end&utm_medium=csdn&utm_campaign=rewrite&utm_content= 。采样请求的调用配额会更充裕,适合长期跑。
最后提醒一点:Sampling 请求里的modelPreferences.hints只是建议,Client 最终选哪个模型由costPriority、speedPriority、intelligencePriority三个参数决定。你可以根据实际场景调整这三个值的权重,让模型选择更符合你的成本和质量要求。