1. 案例目标
本案例展示了如何使用LlamaIndex工作流(Workflow)实现一个带有内联引用的RAG(检索增强生成)查询引擎。主要目标包括:
- 实现一个能够为生成答案提供精确引用的RAG系统
- 展示如何使用工作流构建多步骤的RAG处理流程
- 演示如何将检索到的节点分割为更小的引用块
- 展示如何在生成的答案中嵌入引用标记
- 提供完整的端到端实现,从数据加载到查询响应
2. 技术栈与核心依赖
from llama_index.core.workflow import ( Event, Context, Workflow, StartEvent, StopEvent, step, ) from llama_index.core import SimpleDirectoryReader, VectorStoreIndex from llama_index.llms.openai import OpenAI from llama_index.embeddings.openai import OpenAIEmbedding from llama_index.core.prompts import PromptTemplate from llama_index.core.response_synthesizers import get_response_synthesizer核心依赖包括:
- LlamaIndex工作流框架:用于构建和编排多步骤处理流程
- 向量存储索引:用于文档的向量化存储和检索
- OpenAI模型:包括LLM和嵌入模型,用于生成和向量化
- 响应合成器:用于基于检索到的节点生成最终答案
- 事件系统:用于工作流步骤间的数据传递
3. 环境配置
# 安装依赖 !pip install -U llama-index # 设置OpenAI API密钥 os.environ["OPENAI_API_KEY"] = "sk-..." # 下载示例数据 !mkdir -p 'data/paul_graham/' !wget 'https://raw.githubusercontent.com/run-llama/llama_index/main/docs/examples/data/paul_graham/paul_graham_essay.txt' -O 'data/paul_graham/paul_graham_essay.txt'环境配置包括:
- 安装最新版本的LlamaIndex库
- 设置OpenAI API密钥,确保能够访问GPT模型
- 下载示例数据(Paul Graham的文章)
- 创建数据存储目录
4. 案例实现
4.1 工作流设计
CitationQueryEngine工作流包含以下步骤:
- 索引数据,创建向量索引
- 使用索引和查询检索相关节点
- 为检索到的节点添加引用标记
- 合成最终响应,包含内联引用
4.2 定义事件
为了处理这些步骤,需要定义几个事件:
from llama_index.core.workflow import Event from llama_index.core.schema import NodeWithScore class RetrieverEvent(Event): """检索结果事件""" nodes: list[NodeWithScore] class CreateCitationsEvent(Event): """添加引用事件""" nodes: list[NodeWithScore]4.3 引用提示模板
定义用于生成带引用答案的提示模板:
CITATION_QA_TEMPLATE = PromptTemplate( "Please provide an answer based solely on the provided sources. " "When referencing information from a source, " "cite the appropriate source(s) using their corresponding numbers. " "Every answer should include at least one source citation. " "Only cite a source when you are explicitly referencing it. " "If none of the sources are helpful, you should indicate that. " "For example:\n" "Source 1:\n" "The sky is red in the evening and blue in the morning.\n" "Source 2:\n" "Water is wet when the sky is red.\n" "Query: When is water wet?\n" "Answer: Water will be wet when the sky is red [2], " "which occurs in the evening [1].\n" "Now it's your turn. Below are several numbered sources of information:" "\n------\n" "{context_str}" "\n------\n" "Query: {query_str}\n" "Answer: " )4.4 工作流实现
步骤1: 检索节点
@step async def retrieve( self, ctx: Context, ev: StartEvent ) -> Union[RetrieverEvent, None]: """RAG的入口点,由带有查询的StartEvent触发""" query = ev.get("query") if not query: return None print(f"Query the database with: {query}") # 将查询存储在全局上下文中 await ctx.store.set("query", query) if ev.index is None: print("Index is empty, load some documents before querying!") return None retriever = ev.index.as_retriever(similarity_top_k=2) nodes = retriever.retrieve(query) print(f"Retrieved {len(nodes)} nodes.") return RetrieverEvent(nodes=nodes)步骤2: 创建引用节点
@step async def create_citation_nodes( self, ev: RetrieverEvent ) -> CreateCitationsEvent: """ 修改检索到的节点,为引用创建细粒度源。 接受NodeWithScore对象列表,并将其内容分割为更小的块, 为每个块创建新的NodeWithScore对象。每个新节点都标记为编号源, 允许在查询结果中进行更精确的引用。 Args: nodes (List[NodeWithScore]): 要处理的NodeWithScore对象列表。 Returns: List[NodeWithScore]: 新的NodeWithScore对象列表,其中每个对象 代表原始节点的较小块,标记为源。 """ nodes = ev.nodes new_nodes: List[NodeWithScore] = [] text_splitter = SentenceSplitter( chunk_size=DEFAULT_CITATION_CHUNK_SIZE, chunk_overlap=DEFAULT_CITATION_CHUNK_OVERLAP, ) for node in nodes: text_chunks = text_splitter.split_text( node.node.get_content(metadata_mode=MetadataMode.NONE) ) for text_chunk in text_chunks: text = f"Source {len(new_nodes)+1}:\n{text_chunk}\n" new_node = NodeWithScore( node=TextNode.parse_obj(node.node), score=node.score ) new_node.node.text = text new_nodes.append(new_node) return CreateCitationsEvent(nodes=new_nodes)步骤3: 合成响应
@step async def synthesize( self, ctx: Context, ev: CreateCitationsEvent ) -> StopEvent: """使用检索到的节点返回流式响应""" llm = OpenAI(model="gpt-4o-mini") query = await ctx.store.get("query", default=None) synthesizer = get_response_synthesizer( llm=llm, text_qa_template=CITATION_QA_TEMPLATE, refine_template=CITATION_REFINE_TEMPLATE, response_mode=ResponseMode.COMPACT, use_async=True, ) response = await synthesizer.asynthesize(query, nodes=ev.nodes) return StopEvent(result=response)4.5 创建索引
documents = SimpleDirectoryReader("data/paul_graham").load_data() index = VectorStoreIndex.from_documents( documents=documents, embed_model=OpenAIEmbedding(model_name="text-embedding-3-small"), )4.6 运行工作流
# 创建工作流实例 w = CitationQueryEngineWorkflow() # 运行查询 result = await w.run(query="What information do you have", index=index)4.7 查看引用
# 显示结果 display(Markdown(f"{result}")) # 查看引用源 print(result.source_nodes[0].node.get_text()) print(result.source_nodes[1].node.get_text())5. 案例效果
本案例实现了以下效果:
- 精确引用:生成的答案中每个事实都带有对应的引用标记
- 细粒度源分割:将原始文档分割为更小的块,提供更精确的引用
- 清晰的工作流:将RAG过程分解为明确的步骤,便于理解和维护
- 可验证性:用户可以根据引用标记追溯到原始文档内容
示例输出
The provided sources contain various insights into Paul Graham's experiences and thoughts on programming, writing, and his educational journey. For instance, he reflects on his early experiences with programming on the IBM 1401, where he struggled to create meaningful programs due to the limitations of the technology at the time [2]. He also describes his transition to using microcomputers, which allowed for more interactive programming experiences [3]. Additionally, Graham shares his initial interest in philosophy during college, which he later found less engaging compared to the fields of artificial intelligence and programming [3]. Overall, the sources highlight his evolution as a writer and programmer, as well as his changing academic interests.
关键特性
本案例中的引用系统具有以下特点:
- 每个引用块都有明确的编号,如[1]、[2]等
- 引用块大小可配置,默认为512个字符,重叠20个字符
- 答案中引用的格式遵循学术标准,便于验证
- 引用内容保留原始文档的上下文信息
6. 案例实现思路
本案例的实现思路如下:
- 工作流设计:将RAG过程分解为检索、引用创建和响应合成三个主要步骤
- 事件驱动:使用事件系统在工作流步骤间传递数据和状态
- 引用分割:将检索到的文档分割为更小的块,每个块标记为编号源
- 提示工程:设计专门的提示模板,指导模型在答案中包含引用
- 上下文管理:使用工作流上下文存储查询信息,在步骤间共享
工作流数据流
- StartEvent(查询, 索引) → retrieve步骤
- retrieve步骤 → RetrieverEvent(检索到的节点)
- RetrieverEvent → create_citation_nodes步骤
- create_citation_nodes步骤 → CreateCitationsEvent(带引用的节点)
- CreateCitationsEvent → synthesize步骤
- synthesize步骤 → StopEvent(带引用的响应)
7. 扩展建议
基于本案例,可以考虑以下扩展方向:
- 多模态引用:扩展支持图像、表格等多模态内容的引用
- 引用质量评估:实现引用相关性和准确性的自动评估
- 动态引用调整:根据查询复杂度动态调整引用块大小和数量
- 引用排序:基于相关性对引用进行排序和筛选
- 交互式引用:允许用户点击引用查看完整上下文
- 引用格式定制:支持多种引用格式,如APA、MLA等
- 引用网络分析:分析引用间的关系,构建知识图谱
- 跨文档引用:支持跨多个文档的引用和关联
8. 总结
本案例展示了如何使用LlamaIndex工作流实现一个带有内联引用的RAG查询引擎。案例中的关键技术点包括:
- 工作流编排:使用LlamaIndex工作流框架构建多步骤处理流程
- 引用系统:将检索到的文档分割为带编号的引用块
- 提示工程:设计专门的提示模板,指导模型生成带引用的答案
- 事件驱动:使用事件系统在工作流步骤间传递数据和状态
通过这种方式,生成的答案不仅内容准确,而且每个事实都可以追溯到原始文档,大大提高了RAG系统的可信度和可验证性。这种引用系统特别适用于学术研究、法律分析、医疗诊断等对准确性和可追溯性要求较高的场景。