如何阻止Llama Index向嵌入文档的Metadata添加冗余内容?
问题
使用Llama Index处理文档嵌入时,自定义Metadata能正常存入数据库,但框架会自动添加大量冗余内容(如_node_content、document_id等),即便手动自定义Metadata或尝试pop移除也无效。
自定义Metadata示例:
{"file_path": "path/to/text.txt", "file_name": "myTextFiles.txt", "file_type": "text/plain", "file_size": 3024349, "creation_date": "2024-05-19", "last_modified_date": "2023-11-24"}
自动添加的冗余内容示例:
"_node_content": "{\"id_\": \"e0faec05-8a68-43b2-a2d1-51c307775877\", \"embedding\": null, \"metadata\": {\"file_path\": \"/path/to/textFiles.txt\", \"file_name\": \"paul_graham_essays.txt\", \"file_type\": \"text/plain\", \"file_size\": 3024349, \"creation_date\": \"2024-05-19\", \"last_modified_date\": \"2023-11-24\"}, \"excluded_embed_metadata_keys\": [\"file_name\", \"file_type\", \"file_size\", \"creation_date\", \"last_modified_date\", \"last_accessed_date\"], \"excluded_llm_metadata_keys\": [\"file_name\", \"file_type\", \"file_size\", \"creation_date\", \"last_modified_date\", \"last_accessed_date\"], \"relationships\": {\"1\": {\"node_id\": \"b51cd20a-6dbd-4d1b-b46b-4aa4f4d3d358\", \"node_type\": \"4\", \"metadata\": {\"file_path\": \"file_name\": \"texts.txt\", \"file_type\": \"text/plain\", \"file_size\": 3024349, \"creation_date\": \"2024-05-19\", \"last_modified_date\": \"2023-11-24\"}, \"hash\": \"412d644dc9ebf3d7aab8e41560ad724ffa0bc36922ce428305ddd694c2b41b3a\", \"class_name\": \"RelatedNodeInfo\"}, \"2\": {\"node_id\": \"74b97b26-d3b4-4b84-ac0e-e401922ff5f9\", \"node_type\": \"1\", \"metadata\": {\"file_path\": \"file_name\": \"TEXTS.txt\", \"file_type\": \"text/plain\", \"file_size\": 3024349, \"creation_date\": \"2024-05-19\", \"last_modified_date\": \"2023-11-24\"}, \"hash\": \"708868dc8c11472299c40c4ad44d643e9aea2dbe9c1bef325cc0dd0336d25d19\", \"class_name\": \"RelatedNodeInfo\"}, \"3\": {\"node_id\": \"d82535b2-421f-46ca-9972-df8d1e5a0df6\", \"node_type\": \"1\", \"metadata\": {}, \"hash\": \"e65800e1e75593d6e58717024fbcf523f02a9f7a9e7a5f9ea739c7c5780fb26f\", \"class_name\": \"RelatedNodeInfo\"}}, \"text\": \"\", \"start_char_idx\": 3767, \"end_char_idx\": 8347, \"text_template\": \"{metadata_str}\\n\\n{content}\", \"metadata_template\": \"{key}: {value}\", \"metadata_seperator\": \"\\n\", \"class_name\": \"TextNode\"}", "_node_type": "TextNode", "document_id": "b51cd20a-6dbd-4d1b-b46b-4aa4f4d3d358", "doc_id": "b51cd20a-6dbd-4d1b-b46b-4aa4f4d3d358", "ref_doc_id": "b51cd20a-6dbd-4d1b-b46b-4aa4f4d3d358"}
解决方案
1. 自定义Node Parser后置处理器过滤字段
创建节点解析器时,通过metadata_postprocessor自定义逻辑,只保留需要的Metadata字段:
from llama_index.node_parser import SimpleNodeParser def filter_metadata(node): # 定义需要保留的自定义字段集合 keep_keys = {"file_path", "file_name", "file_type", "file_size", "creation_date", "last_modified_date"} node.metadata = {k: v for k, v in node.metadata.items() if k in keep_keys} return node # 初始化带后置处理器的节点解析器 node_parser = SimpleNodeParser.from_defaults( metadata_postprocessor=filter_metadata ) # 用该解析器处理文档生成节点 nodes = node_parser.get_nodes_from_documents(documents)
2. 配置向量存储排除冗余字段
如果使用向量存储(如Chroma、Pinecone),初始化时通过exclude_metadata_keys指定要排除的系统字段:
from llama_index.vector_stores import ChromaVectorStore vector_store = ChromaVectorStore( chroma_collection=your_collection, exclude_metadata_keys=["_node_content", "_node_type", "document_id", "doc_id", "ref_doc_id"] )
3. 嵌入前手动清理节点Metadata
遍历生成的节点,直接删除冗余字段或替换为自定义Metadata:
for node in nodes: # 删除指定冗余字段 redundant_keys = ["_node_content", "_node_type", "document_id", "doc_id", "ref_doc_id"] for key in redundant_keys: node.metadata.pop(key, None) # 或者直接覆盖为自定义Metadata(确保字段匹配) # node.metadata = { # "file_path": node.metadata.get("file_path"), # "file_name": node.metadata.get("file_name"), # ... # }
内容的提问来源于stack exchange,提问作者John Taylor
相关产品推荐
相关产品推荐

