- 文档
- 教程
- 提示工程
- 大模型
- 人工智能
- RAG
- AI Agent
【免费下载链接】Prompt-Engineering-Guide
🐙 Guides, papers, lessons, notebooks and resources for prompt engineering, context engineering, RAG, and AI Agents.
本文是 Prompt-Engineering-Guide 仓库中「Advanced Prompting」系列指南(guides/prompts-advanced-usage.md)的技术解读与实践手册,系统梳理了零样本提示、少样本提示、思维链(CoT)、零样本 CoT、自洽性(Self-Consistency)、生成知识提示与自动提示工程师(APE)七大进阶技术。读完本文,你将理解每种技术的适用场景、完整提示模板与输出效果,并能结合仓库中的双语 Web 版本(pages/techniques 目录下各语言*.en.mdx文档)与 Notebook 动手验证,构建一套可复制的进阶提示策略。
为什么需要更高级的提示技术
在基础指南(guides/prompts-basic-usage.md)中,我们展示了文本摘要、信息抽取、问答、分类、对话、代码生成与基础推理等任务。随着任务复杂度的上升,仅仅给出指令往往不够——模型在需要多步推理的任务上会频繁出错。本文要解决的核心问题就是:当零样本与少样本提示失效时,如何用更系统的提示策略让 LLM 完成算术、常识与符号推理等复杂任务。
下文将按从简单到高级的顺序,逐一展开这些技术。
零样本提示(Zero-Shot Prompting)
现代 LLM(如 GPT-3.5 Turbo、GPT-4、Claude 3)经过大规模训练并针对指令进行了调优,天然具备**零样本(zero-shot)**执行任务的能力——即提示中不包含任何示例或演示(demonstration),模型直接根据指令完成任务。
最典型的零样本示例是情感分类:
Classify the text into neutral, negative, or positive. Text: I think the vacation is okay. Sentiment:模型输出:
Neutral注意上面的提示没有提供任何"文本—标签"配对样例,模型却能输出正确情感,这正是零样本能力的体现。从仓库的 pages/techniques/zeroshot.en.mdx 可知,这一能力背后依赖两类关键技术:
- 指令微调(instruction tuning):在由指令描述的数据集上对模型进行微调,显著提升零样本学习表现(源自 Wei et al. 2022 所引的指令微调工作);
- 基于人类反馈的强化学习(RLHF):将指令微调进一步对齐到人类偏好,这正是 ChatGPT 类模型的基石。
适用判断:当零样本提示失效(模型给出错误或不稳定结果)时,应优先尝试在提示中补充演示或示例,即进入少样本提示阶段。
少样本提示(Few-Shot Prompting)
少样本提示通过**上下文学习(in-context learning)**在提示中提供若干演示,让演示作为后续输入的条件(conditioning),引导模型生成更符合预期的响应。这是应对复杂任务最直接的增强手段。
1-shot 演示:让模型学会"造词造句"
以下示例来自 Brown et al. 2020(GPT-3 论文),任务是根据定义正确使用一个新造的词:
A "whatpu" is a small, furry animal native to Tanzania. An example of a sentence that uses the word whatpu is: We were traveling in Africa and we saw these very cute whatpus. To do a "farduddle" means to jump up and down really fast. An example of a sentence that uses the word farduddle is:模型输出:
When we won the game, we all started to farduddle in celebration.只给 1 个示例(1-shot),模型就掌握了规律。对更困难的任务,可以实验性地增加演示数量(3-shot、5-shot、10-shot 等),观察性能变化。
设计演示的关键要点(Min et al. 2022)
Min et al. (2022) 的研究给出三条关于演示设计的结论,值得在构建少样本提示时反复对照:
- 标签空间与输入文本分布都很重要——无论单个输入的标签是否正确,"演示所规定的标签空间与输入文本分布"都会显著影响性能;
- 格式本身影响性能——即使使用随机标签,也比完全没有标签好得多;
- 标签采样方式有讲究——从标签的真实分布中采样随机标签(而非均匀分布)会带来额外收益。
实验:随机标签仍然有效
下面的提示故意把 Negative/Positive 标签随机分配给输入:
This is awesome! // Negative This is bad! // Positive Wow that movie was rad! // Positive What a horrible show! //输出:
Negative尽管标签被随机化,模型依然得到了正确答案——格式被保留是关键因素之一。更进一步,较新的 GPT 系列模型对随机格式也表现出更强的鲁棒性,例如:
Positive This is awesome! This is bad! Negative Wow that movie was rad! Positive What a horrible show! --输出依然是:
Negative但需要注意:该结论尚未在更复杂、更多样的任务上得到系统验证,实际应用时应自行测试。
少样本提示的局限
少样本提示在多步推理任务上并不稳定。以"奇数之和是否为偶数"的推理题为例,直接零样本询问时模型会给出错误答案:
The odd numbers in this group add up to an even number: 15, 32, 5, 13, 82, 7, 1. A:模型错误输出:
Yes, the odd numbers in this group add up to 107, which is an even number.补充 4 个演示后(只给出 True/False 结论,不展示推理过程):
The odd numbers in this group add up to an even number: 4, 8, 9, 15, 12, 2, 1. A: The answer is False. The odd numbers in this group add up to an even number: 17, 10, 19, 4, 8, 12, 24. A: The answer is True. The odd numbers in this group add up to an even number: 16, 11, 14, 4, 8, 13, 24. A: The answer is True. The odd numbers in this group add up to an even number: 17, 9, 10, 12, 13, 4, 2. A: The answer is False. The odd numbers in this group add up to an even number: 15, 32, 5, 13, 82, 7, 1. A:模型输出:
The answer is True.答案依然错误(正确应为 False)。这说明:该任务需要多个推理步骤,而仅仅提供"结论型"演示不足以让模型学会推理路径。由此引出的结论是:当零样本与少样本都不够用时,需要把问题拆解成步骤并展示给模型——这正是思维链提示的出发点。
思维链提示(Chain-of-Thought Prompting)
思维链(CoT)由 Wei et al. (2022) 提出,核心思想是通过中间推理步骤激发模型的复杂推理能力,并可将其与少样本提示结合使用,在需要"先推理再回答"的任务上取得更好效果。
继续使用上面的推理题,但这次在每个演示中展示完整的推理步骤:
The odd numbers in this group add up to an even number: 4, 8, 9, 15, 12, 2, 1. A: Adding all the odd numbers (9, 15, 1) gives 25. The answer is False. The odd numbers in this group add up to an even number: 17, 10, 19, 4, 8, 12, 24. A: Adding all the odd numbers (17, 19) gives 36. The answer is True. The odd numbers in this group add up to an even number: 16, 11, 14, 4, 8, 13, 24. A: Adding all the odd numbers (11, 13) gives 24. The answer is True. The odd numbers in this group add up to an even number: 17, 9, 10, 12, 13, 4, 2. A: Adding all the odd numbers (17, 9, 13) gives 39. The answer is False. The odd numbers in this group add up to an even number: 15, 32, 5, 13, 82, 7, 1. A:模型输出:
Adding all the odd numbers (15, 5, 13, 7, 1) gives 41. The answer is False.结果完全正确。更有趣的是,即使只提供 1 个带推理步骤的演示也足够:
The odd numbers in this group add up to an even number: 4, 8, 9, 15, 12, 2, 1. A: Adding all the odd numbers (9, 15, 1) gives 25. The answer is False. The odd numbers in this group add up to an even number: 15, 32, 5, 13, 82, 7, 1. A:模型输出:
Adding all the odd numbers (15, 5, 13, 7, 1) gives 41. The answer is False.需要说明的是,论文作者指出这种能力属于涌现能力(emergent ability),只有在模型规模足够大时才会出现。仓库中完整的 CoT 讲解与演示图见 pages/techniques/cot.en.mdx。
零样本思维链(Zero-Shot CoT)
Kojima et al. (2022) 提出了一种极简的零样本思维链方案:在原提示末尾追加一句 "Let's think step by step",即可触发模型分步推理。这在没有足够示例可用的场景下尤其有价值。
先看不加特殊提示的普通提问:
I went to the market and bought 10 apples. I gave 2 apples to the neighbor and 2 to the repairman. I then went and bought 5 more apples and ate 1. How many apples did I remain with?模型错误输出:
11 apples加上特殊提示后再试:
I went to the market and bought 10 apples. I gave 2 apples to the neighbor and 2 to the repairman. I then went and bought 5 more apples and ate 1. How many apples did I remain with? Let's think step by step.模型输出:
First, you started with 10 apples. You gave away 2 apples to the neighbor and 2 to the repairman, so you had 6 apples left. Then you bought 5 more apples, so now you had 11 apples. Finally, you ate 1 apple, so you would remain with 10 apples.仅仅一行提示就能把错误答案(11)纠正为正确答案(10),且模型完整展示了推理链条。这正是零样本 CoT 的魅力:用最小的提示成本换取可解释的推理过程。
自洽性(Self-Consistency)
自洽性由 Wang et al. (2022) 提出,目标是取代思维链提示中朴素的贪心解码:通过少样本 CoT 采样出多条多样化的推理路径,再基于这些生成结果选出最一致的答案,从而在算术与常识推理任务上进一步提升 CoT 的性能。
先看一个单次解码会出错的例子:
When I was 6 my sister was half my age. Now I’m 70 how old is my sister?模型错误输出:
35使用 Wang et al. 2022 论文(Table 17)中的少样本示例,配合多次采样:
Q: There are 15 trees in the grove. Grove workers will plant trees in the grove today. After they are done, there will be 21 trees. How many trees did the grove workers plant today? A: We start with 15 trees. Later we have 21 trees. The difference must be the number of trees they planted. So, they must have planted 21 - 15 = 6 trees. The answer is 6. Q: If there are 3 cars in the parking lot and 2 more cars arrive, how many cars are in the parking lot? A: There are 3 cars in the parking lot already. 2 more arrive. Now there are 3 + 2 = 5 cars. The answer is 5. Q: Leah had 32 chocolates and her sister had 42. If they ate 35, how many pieces do they have left in total? A: Leah had 32 chocolates and Leah’s sister had 42. That means there were originally 32 + 42 = 74 chocolates. 35 have been eaten. So in total they still have 74 - 35 = 39 chocolates. The answer is 39. Q: Jason had 20 lollipops. He gave Denny some lollipops. Now Jason has 12 lollipops. How many lollipops did Jason give to Denny? A: Jason had 20 lollipops. Since he only has 12 now, he must have given the rest to Denny. The number of lollipops he has given to Denny must have been 20 - 12 = 8 lollipops. The answer is 8. Q: Shawn has five toys. For Christmas, he got two toys each from his mom and dad. How many toys does he have now? A: He has 5 toys. He got 2 from mom, so after that he has 5 + 2 = 7 toys. Then he got 2 more from dad, so in total he has 7 + 2 = 9 toys. The answer is 9. Q: There were nine computers in the server room. Five more computers were installed each day, from monday to thursday. How many computers are now in the server room? A: There are 4 days from monday to thursday. 5 computers were added each day. That means in total 4 * 5 = 20 computers were added. There were 9 computers in the beginning, so now there are 9 + 20 = 29 computers. The answer is 29. Q: Michael had 58 golf balls. On tuesday, he lost 23 golf balls. On wednesday, he lost 2 more. How many golf balls did he have at the end of wednesday? A: Michael initially had 58 balls. He lost 23 on Tuesday, so after that he has 58 - 23 = 35 balls. On Wednesday he lost 2 more so now he has 35 - 2 = 33 balls. The answer is 33. Q: Olivia has $23. She bought five bagels for $3 each. How much money does she have left? A: She bought 5 bagels for $3 each. This means she spent 5 Q: When I was 6 my sister was half my age. Now I’m 70 how old is my sister? A:多次采样得到的输出:
输出 1:
When I was 6 my sister was half my age, so she was 3. Now I am 70, so she is 70 - 3 = 67. The answer is 67.输出 2:
When the narrator was 6, his sister was half his age, which is 3. Now that the narrator is 70, his sister would be 70 - 3 = 67 years old. The answer is 67.输出 3:
When I was 6 my sister was half my age, so she was 3. Now I am 70, so she is 70/2 = 35. The answer is 35.可以看到,前两条路径都得到67,出现多数一致答案,最终即可确定为 67。真实应用中"求最终答案"还需结合投票等步骤(详见原论文),但从工程角度理解其核心:用"多次采样 + 一致性投票"对冲单次贪心解码的偶然错误。仓库中的完整演示见 pages/techniques/consistency.en.mdx。
生成知识提示(Generated Knowledge Prompting)
LLM 的一个常见短板是缺乏世界知识。生成知识提示的思路来自 Liu et al. 2022:在模型做预测之前,先让模型自己生成相关知识,并把知识并入提示中,再基于知识作答。该方法对常识推理类任务尤其有帮助。
第一步:看一个常识错误
Part of golf is trying to get a higher point total than others. Yes or No?模型错误输出:
Yes.这个错误暴露出 LLM 对世界知识的理解局限——高尔夫恰恰是杆数越低越好。
第二步:先生成若干条知识
提示模型为一系列输入生成知识:
Input: Greece is larger than mexico. Knowledge: Greece is approximately 131,957 sq km, while Mexico is approximately 1,964,375 sq km, making Mexico 1,389% larger than Greece. Input: Glasses always fog up. Knowledge: Condensation occurs on eyeglass lenses when water vapor from your sweat, breath, and ambient humidity lands on a cold surface, cools, and then changes into tiny drops of liquid, forming a film that you see as fog. Your lenses will be relatively cool compared to your breath, especially when the outside air is cold. Input: A fish is capable of thinking. Knowledge: Fish are more intelligent than they appear. In many areas, such as memory, their cognitive powers match or exceed those of ’higher’ vertebrates including non-human primates. Fish’s long-term memories help them keep track of complex social relationships. Input: A common effect of smoking lots of cigarettes in one’s lifetime is a higher than normal chance of getting lung cancer. Knowledge: Those who consistently averaged less than one cigarette per day over their lifetime had nine times the risk of dying from lung cancer than never smokers. Among people who smoked between one and 10 cigarettes per day, the risk of dying from lung cancer was nearly 12 times higher than that of never smokers. Input: A rock is the same size as a pebble. Knowledge: A pebble is a clast of rock with a particle size of 4 to 64 millimetres based on the Udden-Wentworth scale of sedimentology. Pebbles are generally considered larger than granules (2 to 4 millimetres diameter) and smaller than cobbles (64 to 256 millimetres diameter). Input: Part of golf is trying to get a higher point total than others. Knowledge:模型为最后一个输入生成两条候选知识:
知识 1:
The objective of golf is to play a set of holes in the least number of strokes. A round of golf typically consists of 18 holes. Each hole is played once in the round on a standard golf course. Each stroke is counted as one point, and the total number of strokes is used to determine the winner of the game.知识 2:
Golf is a precision club-and-ball sport in which competing players (or golfers) use many types of clubs to hit balls into a series of holes on a course using the fewest number of strokes. The goal is to complete the course with the lowest score, which is calculated by adding up the total number of strokes taken on each hole. The player with the lowest score wins the game.第三步:整合知识做出预测
将问题改写为 QA 格式并拼接知识,引导模型给出带解释的答案:
Question: Part of golf is trying to get a higher point total than others. Yes or No? Knowledge: The objective of golf is to play a set of holes in the least number of strokes. A round of golf typically consists of 18 holes. Each hole is played once in the round on a standard golf course. Each stroke is counted as one point, and the total number of strokes is used to determine the winner of the game. Explain and Answer:答案 1(置信度非常高):
No, the objective of golf is not to get a higher point total than others. Rather, the objective is to play a set of holes in the least number of strokes. The total number of strokes is used to determine the winner of the game, not the total number of points.换用知识 2 再试:
Question: Part of golf is trying to get a higher point total than others. Yes or No? Knowledge: Golf is a precision club-and-ball sport in which competing players (or golfers) use many types of clubs to hit balls into a series of holes on a course using the fewest number of strokes. The goal is to complete the course with the lowest score, which is calculated by adding up the total number of strokes taken on each hole. The player with the lowest score wins the game. Explain and Answer:答案 2(置信度明显更低):
Yes, part of golf is trying to get a higher point total than others. Each player tries to complete the course with the lowest score, which is calculated by adding up the total number of strokes taken on each hole. The player with the lowest score wins the game.这个对比非常有意思:不同知识会带来不同结论与置信度,说明知识的质量直接决定预测质量。工程化落地时,通常需要结合多条候选知识的答案与置信度做汇总(详见原论文)。本节的完整示例见仓库 pages/techniques/knowledge.en.mdx。
自动提示工程师(Automatic Prompt Engineer, APE)
Zhou et al. (2022) 提出了 APE 框架,用于自动化地生成与筛选指令:把指令生成问题建模为自然语言综合任务,并作为一个黑盒优化问题,由 LLM 自身生成候选方案并搜索最优指令。
APE 的两阶段流程
- 候选指令生成:由一个 LLM(作为推理模型)接收任务的输出演示(output demonstrations),生成一批指令候选;
- 候选指令评估与选择:用目标模型执行这些指令,基于计算出的评估分数(evaluation scores)选出最合适的指令。
这一"生成—执行—打分—选择"的闭环,把提示工程从纯人工试错提升为可自动化的搜索过程。
APE 的经典成果:发现优于人工的零样本 CoT 提示
APE 发现了一个比人工设计的 "Let's think step by step"(Kojima et al. 2022)更好的零样本 CoT 提示:
"Let's work this out in a step by step way to be sure we have the right answer."
该提示能够触发思维链推理,并在 MultiArith 与 GSM8K 基准上提升性能:
这直接证明了自动提示优化在真实基准上的价值——机器发现的提示可以超越人类专家手工编写的提示。
相关自动优化方向
若对提示自动优化感兴趣,原文档与 pages/techniques/ape.en.mdx 还整理了以下代表性工作:
- AutoPrompt:基于梯度引导搜索,为多样化任务自动创建提示;
- Prefix Tuning:微调的轻量替代方案,为 NLG 任务前置一段可训练连续前缀;
- Prompt Tuning:通过反向传播学习软提示(soft prompts)。
这些方法共同构成了"提示优化"这一重要研究分支,是 APE 之外值得深入的方向。
进阶路线与仓库配套资源
掌握本文七项技术后,可以按以下路径在仓库中继续深化:
- 对照 Web 版本精读:本文内容在仓库中以多语言 MDX 页面组织,英文版依次位于 pages/techniques/zeroshot.en.mdx、pages/techniques/fewshot.en.mdx、pages/techniques/cot.en.mdx、pages/techniques/consistency.en.mdx、pages/techniques/knowledge.en.mdx、pages/techniques/ape.en.mdx,可对照中英文版本加深理解;
- 动手实践:仓库提供了配套 Notebook notebooks/pe-lecture.ipynb,可用于在真实 API 上复现本文各技术的提示模板与输出效果;
- 衔接前后章节:本文是进阶部分,向前承接 guides/prompts-basic-usage.md(基础提示),向后衔接 guides/prompts-applications.md(提示工程应用);本系列全部指南的索引见 guides/README.md;
- 注意适用范围:本文各示例中的"模型是否输出正确结果"会随模型版本与参数设置(如温度、采样次数)变化,尤其是自洽性与生成知识技术对采样数量、评估方式敏感,落地前应基于目标模型自行验证。
小结:如何选择进阶提示策略
根据任务类型,可以将本文技术整理为如下决策顺序:
| 场景 | 推荐技术 | 关键操作 |
|---|---|---|
| 模型能直接完成 | 零样本提示 | 仅提供清晰指令 |
| 零样本失效 | 少样本提示 | 在提示中提供 2~5 条演示 |
| 需要多步推理 | 思维链(CoT) | 演示中展示推理步骤 |
| 缺少示例 | 零样本 CoT | 追加 "Let's think step by step" |
| 推理结果不稳定 | 自洽性 | 多次采样 + 一致性投票 |
| 依赖世界知识 | 生成知识提示 | 先生成知识,再基于知识作答 |
| 提示优化成本高 | 自动提示工程师(APE) | 让 LLM 自动生成并筛选指令 |
从零样本到自动提示工程师,这条进阶路径的底层逻辑始终一致:不是模型本身变强了,而是我们学会用更结构化的方式把推理路径、知识与评价标准暴露给模型。把这七项技术配合 notebooks/pe-lecture.ipynb 反复演练,即可在真实业务中构建出稳定、可解释、可复用的高级提示方案。
- 文档
- 教程
- 提示工程
- 大模型
- 人工智能
- RAG
- AI Agent
【免费下载链接】Prompt-Engineering-Guide
🐙 Guides, papers, lessons, notebooks and resources for prompt engineering, context engineering, RAG, and AI Agents.
相关推荐
ARIS 接入 OpenRouter:用免费或按量模型搭建跨模型审稿后端的完整指南
ARIS 接入 OpenRouter:用免费或按量模型搭建跨模型审稿后端的完整指南 ARIS(Auto Research In Sleep)以 Markdown
AI 技能/插件AI 评测科研人工智能MCP 服务dsh-plugin提示工程进阶:掌握思维链技术的完整指南
提示工程进阶:掌握思维链技术的完整指南 欢迎来到Awesome Prompt Engineering项目中的思维链技术深度解析!思维链(Chain of Tho
提示工程教程UI-TARS桌面版终极指南:如何用AI自然语言控制你的电脑和浏览器?
UI TARS桌面版终极指南:如何用AI自然语言控制你的电脑和浏览器? UI TARS desktop是一个基于视觉语言模型的开源多模态AI助手,让你通过自然语
人工智能大模型AI Agent桌面应用GUI 自动化浏览器控制MCP 服务MCP Clients
创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考