如何修改Haystack教程Notebook为Document添加原始文件名元数据
解决方法
要给每个答案关联原始txt文件,只需在导入文档时把文件名存入Document的meta字段,后续通过答案的document_ids即可反向查询到对应文件名。具体修改步骤如下:
1. 修改文档导入代码,添加文件名到Document的meta
原教程中批量导入文件的逻辑不会自动将文件名存入meta,替换成以下代码,手动遍历每个txt文件并创建带文件名meta的Document:
import os from haystack.document_stores import InMemoryDocumentStore from haystack import Document # 初始化DocumentStore(和原教程逻辑一致) document_store = InMemoryDocumentStore(use_bm25=True) # 遍历目标文件夹中的txt文件,逐个生成带文件名meta的Document data_dir = "data/tutorial1" # 替换为你实际的文件存储目录 docs = [] for filename in os.listdir(data_dir): if filename.endswith(".txt"): file_path = os.path.join(data_dir, filename) with open(file_path, "r", encoding="utf-8") as f: text_content = f.read() # 创建Document对象,将文件名存入meta字段 doc = Document(content=text_content, meta={"filename": filename}) docs.append(doc) # 将带meta信息的文档写入DocumentStore document_store.write_documents(docs)
2. 修改答案输出代码,关联原始文件名
运行QA pipeline得到预测结果后,遍历每个答案,通过document_ids从DocumentStore中取出对应文档,再提取meta里的文件名:
from pprint import pprint # 运行pipeline获取答案(和原教程逻辑一致) prediction = pipeline.run( query="你的问题内容", params={"Retriever": {"top_k": 10}, "Reader": {"top_k": 5}} ) # 遍历答案,打印内容及对应来源文件 for answer in prediction["answers"]: # 通过document_id获取对应的Document对象 target_doc = document_store.get_document_by_id(answer.document_ids[0]) # 提取meta中的文件名 source_file = target_doc.meta["filename"] # 自定义输出格式 print(f"答案:{answer.answer}") print(f"来源文件:{source_file}") print(f"置信度:{answer.score:.2f}\n")
如果想用pprint直接查看包含文件名的完整答案信息,也可以给每个答案临时添加来源文件名字段:
# 给每个答案追加来源文件名 for answer in prediction["answers"]: target_doc = document_store.get_document_by_id(answer.document_ids[0]) answer.meta["source_filename"] = target_doc.meta["filename"] # 打印包含文件名的答案详情 pprint(prediction["answers"])
内容的提问来源于stack exchange,提问作者ekohrt
相关产品推荐
相关产品推荐

