如何在Gradio聊天界面展示Langchain源文档元数据
解决方案
要在Gradio聊天界面中展示RAG检索到的PDF源文件信息,只需修改Gradio的交互逻辑,提取对话链返回结果里的元数据并整合到回复内容中即可,具体实现如下:
1. 修改Gradio用户交互函数
在user函数中,从response["source_documents"]里提取PDF路径和页码,将这些信息附加到回答末尾:
def user(user_message, history): # 获取对话链响应 response = conv_chain({"question": user_message, "chat_history": history}) # 提取源文档元数据 source_details = [] for doc in response["source_documents"]: doc_path = doc.metadata["source"] page_num = doc.metadata["page"] # 生成带页码的源信息,本地场景可转为可点击的文件链接 display_text = f"- {os.path.basename(doc_path)} (第{page_num}页)" # 若要本地可点击,替换为:display_text = f"- [{os.path.basename(doc_path)}](file://{os.path.abspath(doc_path)}) (第{page_num}页)" source_details.append(display_text) # 整合回答与源信息 full_reply = f"{response['answer']}\n\n**参考来源:**\n{'\n'.join(source_details)}" # 更新聊天历史 history.append((user_message, full_reply)) return gr.update(value=""), history
2. 配置Gradio允许访问PDF路径
在启动Gradio时,将PDF所在文件夹添加到allowed_paths,避免本地文件访问被拦截:
if __name__ == "__main__": demo.queue(max_size=10) demo.launch( share=False, debug=True, server_name="xxxxxx", server_port=xxx, allowed_paths=["images/xxxx", "pdfs-folder"] # 新增PDF存放文件夹 )
3. 可选:让LLM主动提及来源(按需选择)
如果希望LLM生成的回答中自然包含来源信息,可以修改Prompt模板,明确要求模型引用源文档:
prompt_template: str = """<|system|> You are a helpful, respectful and honest assistant. Use the following pieces of context to answer the question at the end. Always respond to questions using the English language unless asked to do so otherwise. Your answers should not include any harmful, unethical, racist, sexist, toxic, dangerous, or illegal content. Please ensure that your responses are socially unbiased and positive in nature. If a question does not make any sense, or is not factually coherent, explain why instead of answering something incorrectly. If you don't know the answer to a question, please don't share false information. After answering the question, list the source documents you used in this format: "- [file name] (page X)". </s> {context} {chat_history} <|user|> {question} </s> <|assistant|>""" PROMPT = PromptTemplate.from_template(template=prompt_template)
这种方式依赖LLM对元数据的解析能力,相比直接提取元数据附加的方式,稳定性稍弱,按需选择即可。
内容的提问来源于stack exchange,提问作者texasdave
相关产品推荐
相关产品推荐

