You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Llama Index引擎查询后返回检索节点的源文档?

解决Llama Index中获取检索节点对应原始文档的方法

要获取检索节点对应的原始文档并统计引用频率,核心是利用Llama Index中SourceNode的metadata属性——每个拆分后的节点都会继承原始文档的元数据信息,以下是具体实现步骤:

1. 确保原始文档加载时包含唯一标识

加载文档时,为每个文档添加可识别的元数据(如文件名、自定义文档ID),方便后续关联:

from llama_index import SimpleDirectoryLoader

def attach_doc_metadata(doc):
    # 添加文件名作为标识(或自定义document_id)
    doc.metadata["file_name"] = doc.metadata.get("file_name")
    doc.metadata["document_id"] = doc.id_  # 使用Llama Index自动生成的文档ID
    return doc

# 加载文档时应用元数据处理函数
loader = SimpleDirectoryLoader("./your_docs_dir", metadata_fn=attach_doc_metadata)
documents = loader.load_data()

2. 从检索结果中提取原始文档信息

修改你的查询循环代码,直接从response.source_nodes中提取每个节点对应的原始文档标识:

source_nodes = []
referenced_docs = []  # 存储每个查询引用的原始文档列表

with tru_recorder as recording:
    for question in eval_questions:
        response = sentence_window_engine_1.query(question)
        source_nodes.append(response.source_nodes)
        
        # 提取每个节点对应的原始文档标识(根据你设置的元数据字段调整)
        current_docs = [
            node.metadata.get("file_name") or node.metadata.get("document_id") 
            for node in response.source_nodes
        ]
        referenced_docs.append(current_docs)

3. 统计文档引用频率

收集所有查询的引用文档后,用计数器统计每个文档的被引用次数:

from collections import Counter

# 扁平化所有引用的文档列表
all_referenced_docs = [doc for sublist in referenced_docs for doc in sublist]
# 统计频率
doc_reference_count = Counter(all_referenced_docs)

# 按引用次数降序排列
sorted_doc_counts = sorted(doc_reference_count.items(), key=lambda x: x[1], reverse=True)

# 打印结果
print("文档引用频率统计:")
for doc, count in sorted_doc_counts:
    print(f"{doc}: {count} 次")

4. 与Trulens Records关联分析

如果需要将文档引用信息和Trulens的评估数据合并,可以将referenced_docs与records DataFrame合并:

import pandas as pd

# 创建文档引用的DataFrame
doc_ref_df = pd.DataFrame({
    "eval_question": eval_questions,
    "referenced_documents": referenced_docs
})

# 合并到Trulens的records(假设records按查询顺序排列)
merged_records = pd.concat([records, doc_ref_df], axis=1)

# 保存或分析合并后的数据
merged_records.to_csv("rag_evaluation_with_docs.csv", index=False)

关键说明

  • Llama Index的每个拆分节点(如句子窗口生成的节点)都会继承原始文档的metadata,因此可以直接从节点中获取原始文档的标识信息。
  • 如果你的文档加载方式不同(如非本地文件),只需调整metadata的提取字段即可,确保每个节点能关联到唯一的原始文档。

内容的提问来源于stack exchange,提问作者JP1990

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 18:06:17