- 大模型
- 基础模型
- 人工智能
- 深度学习
- NLP
【免费下载链接】DeepSeek-V3.1-Base
DeepSeek-V3.1 是一款支持思考模式与非思考模式的混合模型
导读
本文以 DeepSeek-V3.1-Base 仓库(DeepSeek-V3.1-Base)的 README.md 为主线,系统讲解这款 671B 参数混合模型的三大核心技术能力:通过切换聊天模板在同一模型上实现思考/非思考双模式、专门优化的工具调用(ToolCall)与 Agent 能力,以及 128K 长上下文的两阶段扩展方法与 UE8M0 FP8 量化格式。文章将结合仓库内的 config.json、modeling_deepseek.py、tokenizer_config.json 与 assets/chat_template.jinja 等源码,完整还原官方 Prompt 模板、Python 调用示例与本地运行建议,使读者能够直接复现思考模式切换、ToolCall 格式拼装与 Agent 轨迹构建的完整流程。
一、模型概述与核心升级点
DeepSeek-V3.1 是 DeepSeek 系列的混合推理模型(Hybrid Model),其最大特点是在同一个模型权重上同时支持思考模式(Thinking Mode)与思考模式(Non-Thinking Mode),用户不需要加载两套参数,只需改变 Chat Template即可切换两种推理风格。相对前代版本,官方在 README.md 中总结了三点改进:
- 混合思考模式:一个模型通过更换聊天模板同时支持思考与非思考两种模式;
- 更智能的工具调用:经过后训练(post-training)优化,模型在工具使用与 Agent 任务上的表现显著提升;
- 更高思考效率:DeepSeek-V3.1-Think 在回答质量上与 DeepSeek-R1-0528 相当,但响应速度更快。
在训练来源上,DeepSeek-V3.1 基于 DeepSeek-V3.1-Base 进行后训练,而 Base 模型又是在原始 V3 Base 检查点之上,采用两阶段长上下文扩展方法构建的——收集了更多长文档数据,并将 32K 扩展阶段训练量提升 10 倍至630B tokens,128K 扩展阶段提升 3.3 倍至209B tokens。此外,DeepSeek-V3.1 在模型权重与激活(activations)上统一采用UE8M0 FP8 scale data format,以确保与微缩放数据格式(microscaling data formats)兼容。
1.1 模型规格一览
官方 Model Downloads 表格给出了两个可下载变体的关键指标,本仓库即为其中的DeepSeek-V3.1-Base:
| 模型 | 总参数量 | 激活参数量 | 上下文长度 | 下载渠道 |
|---|---|---|---|---|
| DeepSeek-V3.1-Base | 671B | 37B | 128K | HuggingFace / ModelScope |
| DeepSeek-V3.1 | 671B | 37B | 128K | HuggingFace / ModelScope |
671B 总参数量中每次前向只激活 37B 参数,这正是其 MoE(Mixture-of-Experts)架构的直接体现。本仓库的 config.json 提供了与表格对应的架构级佐证:n_routed_experts = 256(路由专家总数)、n_shared_experts = 1(共享专家数)、num_experts_per_tok = 8(每个 token 激活 8 个专家)、topk_group = 4与n_group = 8(分组路由);hidden_size = 7168、num_hidden_layers = 61、num_nextn_predict_layers = 1,而max_position_embeddings = 163840与rope_scaling.type = "yarn"则对应 128K 上下文与 YaRN 外推策略。
1.2 权重结构与长上下文配置
从 model.safetensors.index.json 的weight_map可以看到(共 91991 个张量条目),路由专家的gate_proj / up_proj / down_proj各有 30332 个,与 61 层中每层 256 个路由专家、仅前 3 层为 dense 层(first_k_dense_replace = 3)的结构吻合。此外存在model.layers.61.shared_head.norm.weight与model.layers.61.shared_head.head.weight,对应num_nextn_predict_layers = 1的Multi-Token Prediction (MTP)头,在 modeling_deepseek.py 中以DeepseekV3ForCausalLM之外的附加预测模块实现。
config.json中的quantization_config块进一步确认了 README 强调的 UE8M0 FP8 细节:
"quantization_config": { "activation_scheme": "dynamic", "fmt": "e4m3", "quant_method": "fp8", "weight_block_size": [128, 128], "scale_fmt": "ue8m0" }其中scale_fmt: "ue8m0"即 README 中 UE8M0 缩放格式的配置化体现,权重按 128×128 块进行 FP8 量化,激活采用动态 e4m3 方案。
二、Chat Template:一套模板,两种模式
DeepSeek-V3.1 的聊天模板完整定义在仓库的 tokenizer_config.json 的chat_template字段中,其可读版本为 assets/chat_template.jinja。模板暴露了两个关键控制参数:
thinking(布尔值):True时在生成前缀末尾追加<think>,False时追加</think>;add_generation_prompt(布尔值):是否在消息序列末尾拼接<|Assistant|>与对应的思考开关 token。
模板内部使用 Jinja 命名空间ns维护状态机(is_first / is_tool / system_prompt / is_first_sp / is_last_user),将 system 消息合并为单一system_prompt,并逐条处理 user / assistant / tool 三类消息。下表总结了 README 描述的四类 Prompt 形态:
| 模式 | 轮次 | 完整前缀格式 |
|---|---|---|
| Non-Thinking | 首轮 | <|begin▁of▁sentence|>{system prompt}<|User|>{query}<|Assistant|></think> |
| Non-Thinking | 多轮 | 上下文为各轮拼接,前缀为<|User|>{query}<|Assistant|></think> |
| Thinking | 首轮 | <|begin▁of▁sentence|>{system prompt}<|User|>{query}<|Assistant|><think> |
| Thinking | 多轮 | 前缀为<|User|>{query}<|Assistant|><think> |
2.1 Non-Thinking 模式详解
非思考模式下,模型直接给出答案。与 DeepSeek V3 相比,V3.1 引入了一个额外的</think>token——即使是首轮非思考对话,也会在<|Assistant|>之后显式输出</think>,这一细节在 README 中被特别强调。
多轮场景中,历史上下文的组装规则为:
<|begin▁of▁sentence|>{system prompt}<|User|>{query}<|Assistant|></think>{response}<|end▁of▁sentence|>...<|User|>{query}<|Assistant|></think>{response}<|end▁of▁sentence|>将上述上下文与当前轮前缀拼接,即可得到正确完整的 Prompt。
2.2 Thinking 模式详解
思考模式的首轮前缀为<|begin▁of▁sentence|>{system prompt}<|User|>{query}<|Assistant|><think>,与 DeepSeek-R1 的思路一致:模型先产出隐藏的思考过程(位于<think>与</think>之间),再输出最终答案。多轮思考模板与非思考多轮模板完全相同——也就是说,每一轮上下文中都保留</think>,但最后一轮的思考 token 会被丢弃(模型只从<think>开始续写)。这条规则在 Jinja 模板中体现为:处理历史 assistant 消息时先将content中</think>之前的部分剥离,仅保留其后的答案内容并追加<|end▁of▁sentence|>;只有对当前轮前缀才按thinking参数决定追加<think>还是</think>。
2.3 一个模型同时实现两种模式的原因
从工程角度理解,同一权重之所以能通过模板切换两种模式,是因为思考行为由输入侧的触发 token(<think>与</think>)驱动,而非由参数集区分。模板负责把thinking布尔值翻译成前缀 token,模型在推理时仅需延续前缀即可。这也意味着部署方不需要准备两套模型服务,只需在同一套服务上按请求传入不同模板参数。
三、ToolCall:非思考模式下的工具调用协议
DeepSeek-V3.1 在非思考模式下原生支持 ToolCall。其完整格式为:
<|begin▁of▁sentence|>{system prompt}\n\n{tool_description}<|User|>{query}<|Assistant|></think>其中tool_description的结构由 README 明确定义为:
## Tools You have access to the following tools: ### {tool_name1} Description: {description} Parameters: {json.dumps(parameters)} IMPORTANT: ALWAYS adhere to this exact format for tool use: <|tool▁calls▁begin|><|tool▁call▁begin|>tool_call_name<|tool▁sep|>tool_call_arguments<|tool▁call▁end|>{additional_tool_calls}<|tool▁calls▁end|> Where: - `tool_call_name` must be an exact match to one of the available tools - `tool_call_arguments` must be valid JSON that strictly follows the tool's Parameters Schema - For multiple tool calls, chain them directly without separators or spaces3.1 格式要点解析
tool_call_name必须是可用工具名中精确匹配的一项;tool_call_arguments必须是严格符合工具 Parameters Schema 的合法 JSON;- 多工具调用时,多个
<|tool▁call▁begin|>...<|tool▁call▁end|>直接首尾相连,不加任何分隔符与空格,整体包在<|tool▁calls▁begin|>与<|tool▁calls▁end|>之间。
3.2 模板中的 ToolCall 处理逻辑
Jinja 模板对 ToolCall 消息的处理与上述协议一一对应:
- assistant 消息若携带
tool_calls字段,模板在上一轮 user 消息后先输出<|Assistant|></think>(保持非思考前缀),随后逐个拼接<|tool▁call▁begin|>name<|tool▁sep|>args<|tool▁call▁end|>,最后统一闭合为<|tool▁calls▁end|><|end▁of▁sentence|>; - 工具执行结果消息(role 为
tool)被包装为<|tool▁output▁begin|>{content}<|tool▁output▁end|>,作为下一轮上下文的一部分; - 首条工具调用(
is_first为 false 时)允许在tool_calls之前附带一段自由文本message['content'](模型通常用一行说明即将调用什么工具),其后继调用则不再重复附带。
四、Code-Agent 与 Search-Agent:Agent 轨迹范例
4.1 Code-Agent(代码智能体)
README 指出:模型支持多种代码智能体框架,开发者可完全依据上文 ToolCall 格式自行搭建自己的代码 Agent,仓库中提供了可直接参考的完整交互轨迹示例 assets/code_agent_trajectory.html。从该 HTML 可见其完整调用链:
- system 提示词将模型设定为 “helpful software engineer assistant”,并提供
bash工具(参数为一个 JSON Schema,描述 command 字段); - 模型在多轮交互中反复使用
<|tool▁calls▁begin|><|tool▁call▁begin|>bash<|tool▁sep|>{...}<|tool▁call▁end|><|tool▁calls▁end|>发起命令执行,工具结果以<|tool▁output▁begin|>/<|tool▁output▁end|>包裹返回; - 整个过程与 README 给出的
tool_description模板完全一致,可作为自定义代码 Agent 的参照实现。
4.2 Search-Agent(搜索智能体)
对于需要访问外部或实时信息的复杂问题,DeepSeek-V3.1 支持在思考模式下通过多轮工具调用完成搜索 Agent 任务——模型可利用用户提供的搜索工具。README 提供了两份详细模板:assets/search_tool_trajectory.html 与 assets/search_python_tool_trajectory.html。
从 assets/search_tool_trajectory.html 的第 36 行开始,可以看到一个完整的搜索轨迹样例(原始问题为“某全球旗舰店被称为博物馆的品牌,2025 年 5 月初与一家中国度假酒店集团联名的宣传片艺术指导是谁”):
- system 提示词注入当前日期(
The current date is 2025-08-16, Saturday)、语言一致性约束(必须以用户问题同语言回复)、引用格式约束(引用网页时必须使用[citation:x])以及search工具的 JSON Schema; - 模型在
<|Assistant|></think>后输出分析过程,然后发起<|tool▁calls▁begin|><|tool▁call▁begin|>search<|tool▁sep|>{"questions": "..."}<|tool▁call▁end|><|tool▁calls▁end|>调用; - 工具结果以
[webpage 0 begin]...的网页片段块返回,并包裹在<|tool▁output▁begin|>/<|tool▁output▁end|>中; - 模型基于多轮搜索结果逐步推理,最终给出带
[citation:x]引用的结论,并以<|search▁end|>收束整个搜索回合。
关键观察点:
- 搜索 Agent 在思考模式下运行,因此每一轮都保留
<think>语境(历史回合剥离思考内容,但保留</think>); - 搜索结果被组织为
[webpage N begin] / [webpage N end]的结构化片段,方便模型在海量结果中定位证据; - 若同时提供 Python 能力(assets/search_python_tool_trajectory.html),模型还能在搜索之外执行脚本、组合两种工具完成更深度的信息处理。
这两份 HTML 轨迹与 Code-Agent 轨迹共同构成了官方认可的 Agent 构建范式,开发者可直接复制其中的 system 提示词、工具描述与消息顺序来驱动自己的 Agent 应用。
五、Evaluation:官方基准结果
README 给出了 DeepSeek V3.1(含 NonThinking 与 Thinking 两个变体)相对 DeepSeek V3 0324、DeepSeek R1 0528 的官方评测数据,摘录如下:
| 类别 | 基准(指标) | V3.1-NonThinking | V3 0324 | V3.1-Thinking | R1 0528 |
|---|---|---|---|---|---|
| General | MMLU-Redux (EM) | 91.8 | 90.5 | 93.7 | 93.4 |
| General | MMLU-Pro (EM) | 83.7 | 81.2 | 84.8 | 85.0 |
| General | GPQA-Diamond (Pass@1) | 74.9 | 68.4 | 80.1 | 81.0 |
| General | Humanity's Last Exam (Pass@1) | - | - | 15.9 | 17.7 |
| Search Agent | BrowseComp | - | - | 30.0 | 8.9 |
| Search Agent | BrowseComp_zh | - | - | 49.2 | 35.7 |
| Search Agent | HLE (Python + Search) | - | - | 29.8 | 24.8 |
| Search Agent | SimpleQA | - | - | 93.4 | 92.3 |
| Code | LiveCodeBench (2408-2505) (Pass@1) | 56.4 | 43.0 | 74.8 | 73.3 |
| Code | Codeforces-Div1 (Rating) | - | - | 2091 | 1930 |
| Code | Aider-Polyglot (Acc.) | 68.4 | 55.1 | 76.3 | 71.6 |
| Code Agent | SWE Verified (Agent mode) | 66.0 | 45.4 | - | 44.6 |
| Code Agent | SWE-bench Multilingual (Agent mode) | 54.5 | 29.3 | - | 30.5 |
| Code Agent | Terminal-bench (Terminus 1 framework) | 31.3 | 13.3 | - | 5.7 |
| Math | AIME 2024 (Pass@1) | 66.3 | 59.4 | 93.1 | 91.4 |
| Math | AIME 2025 (Pass@1) | 49.8 | 51.3 | 88.4 | 87.5 |
| Math | HMMT 2025 (Pass@1) | 33.5 | 29.2 | 84.2 | 79.4 |
需要说明的评测条件(README 原注):
- 搜索 Agent 使用官方内部搜索框架评测:商业搜索 API + 网页过滤器 + 128K 上下文窗口;R1-0528 的搜索 Agent 结果则基于预定义工作流评测;
- SWE-bench 使用官方内部代码 Agent 框架评测;
- HLE 使用纯文本子集评测。
从数据可读出两条主线:一是思考模式整体优于非思考模式(尤其在数学、代码、GPQA 等推理密集型任务上);二是V3.1 的搜索与代码 Agent 能力相对前代与 R1 均有明显提升(如 BrowseComp 30.0 vs 8.9、SWE Verified 66.0 vs 45.4)。这些数据与 README“后训练优化显著提升工具使用与 Agent 任务表现”的表述相互印证。
六、Usage Example:完整的 Python 调用
README 提供了可直接运行的transformers调用示例。核心在于:通过apply_chat_template的thinking参数控制思考模式。
import transformers tokenizer = transformers.AutoTokenizer.from_pretrained("deepseek-ai/DeepSeek-V3.1") messages = [ {"role": "system", "content": "You are a helpful assistant"}, {"role": "user", "content": "Who are you?"}, {"role": "assistant", "content": "<think>Hmm</think>I am DeepSeek"}, {"role": "user", "content": "1+1=?"} ] tokenizer.apply_chat_template(messages, tokenize=False, thinking=True, add_generation_prompt=True) # '<|begin▁of▁sentence|>You are a helpful assistant<|User|>Who are you?<|Assistant|></think>I am DeepSeek<|end▁of▁sentence|><|User|>1+1=?<|Assistant|><think>' tokenizer.apply_chat_template(messages, tokenize=False, thinking=False, add_generation_prompt=True) # '<|begin▁of▁sentence|>You are a helpful assistant<|User|>Who are you?<|Assistant|></think>I am DeepSeek<|end▁of▁sentence|><|User|>1+1=?<|Assistant|></think>'6.1 调用要点逐条拆解
thinking=True:在最终生成前缀处追加<think>,模型将从思考 token 开始续写;thinking=False:追加</think>,模型直接输出答案;- 历史 assistant 消息中的思考内容(
<think>Hmm</think>)会被模板剥离,仅保留I am DeepSeek及其后的<|end▁of▁sentence|>,与 README“历史回合剥离思考 token、保留</think>”的规则一致; add_generation_prompt=True是开启生成前缀的必要条件,否则模板只输出历史部分。
该模板的默认行为(当thinking未显式给出时默认为False)与特殊键处理均在 assets/chat_template.jinja 首部有明确实现。注意:示例中的deepseek-ai/DeepSeek-V3.1为官方在线仓库标识,本仓库对应的是其Base检查点(DeepSeek-V3.1-Base),同样遵循上述模板协议。
七、How to Run Locally:本地运行与使用建议
README 明确说明:DeepSeek-V3.1 的模型结构与 DeepSeek-V3 相同,因此本地运行方式可直接参考 DeepSeek-V3 仓库的说明。对本仓库(DeepSeek-V3.1-Base)而言,模型权重以 163 个.safetensors分片的形式存放(model-00001-of-000163.safetensors 至 model-00163-of-000163.safetensors),配套 model.safetensors.index.json 索引文件,配合transformers(仓库内transformers_version标注为 4.44.2)即可加载:
from transformers import AutoTokenizer, AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("./", torch_dtype="auto", device_map="auto")7.1 两条关键使用建议(务必遵守)
README 给出了两条硬性要求,它们直接关系到推理精度与正确性:
mlp.gate.e_score_correction_bias参数必须用 FP32 精度加载与计算。从源码 modeling_deepseek.py 可见,当topk_method == "noaux_tc"时,MoEGate会注册一个形状为(n_routed_experts,)的e_score_correction_bias偏置参数(即 256 个数值)。在 forward 中的 noaux_tc 分支,该偏置被逐 token 加到 sigmoid 门控分数上以修正专家选择(scores_for_choice = scores.view(bsz * seq_len, -1) + self.e_score_correction_bias.unsqueeze(0))。同时源码在门控打分处强制将hidden_states与weight转换为torch.float32参与计算——这从实现侧印证了 README 要求该偏置保持 FP32 的意图:任何低精度化都可能导致专家路由偏差,进而影响整体生成质量。确保 FP8 模型权重与激活采用 UE8M0 scale format。如前文所述,这一要求在 config.json 的
quantization_config(scale_fmt: "ue8m0")与 README 的 DeepGEMM 指引中均有体现。UE8M0 属于微缩放格式(microscaling data formats)家族,其缩放因子以 8 位无符号指数(E8M0)形式存放,配套的块级缩放(weight_block_size: [128, 128])与动态激活量化(activation_scheme: "dynamic")共同保证权重/激活的低精度表示在数值尺度上对齐。
此外,generation_config.json 提供了推荐的采样默认值(do_sample: true、temperature: 0.6、top_p: 0.95),加载模型后可沿用这些超参获得与官方一致的生成行为。
八、仓库源码速览:模型结构的印证
为了让读者对上述内容有源码级把握,这里补充本仓库关键文件的速览:
- modeling_deepseek.py:约 1848 行的完整 Transformers 实现,核心类包括
DeepseekV3Model(Decoder 主干)、DeepseekV3ForCausalLM(生成入口,含lm_head)、DeepseekV3ForSequenceClassification(序列分类头)、MoEGate(分组专家路由门控,实现noaux_tctop-k 选择与 sigmoid 打分)、DeepseekV3MoE(路由 + 共享专家混合模块)、DeepseekV3Attention(MLA 注意力,支持 FlashAttention 2 加速,见is_flash_attn_2_available相关导入与_flash_attention_forward),以及 YaRN 旋转位置编码相关函数; - configuration_deepseek.py:
DeepseekV3Config配置类,其__init__签名完整罗列了上文涉及的全部超参(n_routed_experts、n_group、topk_group、num_experts_per_tok、routed_scaling_factor、scoring_func="sigmoid"、topk_method="noaux_tc"等),与 config.json 一一对应; - tokenizer_config.json 与 assets/chat_template.jinja:聊天模板的压缩版与可读版,
LlamaTokenizerFast分词器,特殊 token 为<|begin▁of▁sentence|>(BOS)、<|end▁of▁sentence|>(EOS/PAD); - assets/code_agent_trajectory.html、assets/search_tool_trajectory.html、assets/search_python_tool_trajectory.html:三份 Agent 完整轨迹样例,是构建 Code-Agent 与 Search-Agent 的最佳参考。
九、总结
DeepSeek-V3.1-Base 仓库呈现了一整套“混合推理 + 工具调用”的完整技术栈:
- 模式切换:
<think>/</think>两个 token +thinking模板参数,实现单一权重上的双模式推理; - 工具协议:
<|tool▁calls▁begin|>...<|tool▁calls▁end|>包裹、<|tool▁sep|>分隔名称与 JSON 参数的标准化 ToolCall 格式,支持多工具并行调用; - Agent 范式:以思考模式承载 Search-Agent、以非思考模式承载 Code-Agent,两份官方轨迹 HTML 即为可复用的模板;
- 精度纪律:FP32 的
e_score_correction_bias与 UE8M0 FP8 量化格式共同保障 671B 模型在低精度推理下的数值正确性。
无论你是要接入 ToolCall 协议、搭建搜索/代码 Agent,还是部署 128K 长上下文推理服务,都可以从本仓库的模板与源码出发,快速完成方案落地。
- 大模型
- 基础模型
- 人工智能
- 深度学习
- NLP
【免费下载链接】DeepSeek-V3.1-Base
DeepSeek-V3.1 是一款支持思考模式与非思考模式的混合模型
创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考