LangChain ConversationalRetrievalQAChain流式异常:仅获完整响应而非单个Token
LangChain ConversationalRetrievalQAChain 无法实现Token级流式处理
问题背景
基于LangChain开发项目,使用ConversationalRetrievalQAChain结合OpenAI,目标是实现Token级增量处理——每个Token生成时就执行自定义格式化与情感分析,但目前只能在流程结束后收到完整响应,无法获取单个Token。
代码示例
const model = new OpenAI({ ... // 相关配置如temperature、modelName等 streaming: true, callbacks: [ { handleLLMNewToken(token) { console.log("Expected to receive individual tokens here, but only getting full response object."); }, }, ], }); const chain = ConversationalRetrievalQAChain.fromLLM( model, vectorstore.asRetriever({ k: 6 }), ... // 相关链配置如qaTemplate和questionGeneratorTemplate ); const responseStream = await chain.stream({ question: "Sample question", chat_history: "Previous conversation..." }); for await (const data of responseStream) { console.log("Expected to iterate over individual tokens, but only receiving full response object."); }
预期行为
handleLLMNewToken回调能逐个记录生成的Tokenfor await循环可遍历单个Token,实现即时处理
原预期描述:I expect the handleLLMNewToken callback to log each generated token individually, and the for await loop to iterate over these individual tokens. This would allow me to process each token as it's written.
实际行为
handleLLMNewToken回调仅记录完整响应对象for await循环仅遍历到单个完整响应对象,而非单个Token
解决方案
1. 使用专门的流式版本链
普通ConversationalRetrievalQAChain的流式能力有限,LangChain提供了StreamingConversationalRetrievalQAChain,原生支持Token级流式输出,替换后即可实现需求:
import { StreamingConversationalRetrievalQAChain } from "langchain/chains"; // 模型配置保持streaming: true不变 const model = new OpenAI({ temperature: 0, modelName: "gpt-3.5-turbo", streaming: true, callbacks: [ { handleLLMNewToken(token) { console.log("Received token:", token); // 现在会逐个打印Token }, }, ], }); const chain = StreamingConversationalRetrievalQAChain.fromLLM( model, vectorstore.asRetriever({ k: 6 }), { returnSourceDocuments: false, // 根据需求调整是否返回源文档 } ); const responseStream = await chain.stream({ question: "Sample question", chat_history: "Previous conversation...", }); // 遍历单个Token for await (const chunk of responseStream) { console.log("Stream chunk:", chunk); // 此处chunk为单个Token或小片段 }
2. 确认模型与配置兼容性
- 确保使用的OpenAI模型(如gpt-3.5-turbo、gpt-4)支持流式输出,旧模型(如text-davinci-003)的流式逻辑可能存在差异
- 检查链的配置,避免自定义模板或参数覆盖流式输出逻辑
3. 验证回调注册
确保callbacks数组中的handleLLMNewToken方法未被其他配置覆盖,链的初始化过程中没有关闭流式传递。
内容的提问来源于stack exchange,提问作者jyorge maccaline
相关产品推荐
相关产品推荐

