使用llama-index向Weaviate加载Nodes时遇WindowsPath序列化错误
解决Weaviate向量存储导入时的WindowsPath JSON序列化错误
问题场景
我在跟随llama-index的《Building Data Ingestion From Scratch》教程实现数据导入,原教程使用Pinecone的代码如下:
from llama_index.vector_stores import PineconeVectorStore vector_store = PineconeVectorStore(pinecone_index=pinecone_index) vector_store.add(nodes)
我将其替换为Weaviate的实现:
from llama_index.vector_stores import WeaviateVectorStore # 构建向量存储 vector_store = WeaviateVectorStore(weaviate_client = client, index_name="SBCZoning") vector_store.add(nodes)
运行时触发TypeError,提示WindowsPath类型对象无法JSON序列化,完整报错堆栈:
TypeError Traceback (most recent call last) Input In [73], in <cell line: 6>() 1 # nodes_to_parse = SimpleNodeParser.get_nodes_from_documents(nodes) 2 # nodes_to_parse = parser.get_nodes_from_documents(nodes) 3 4 # construct vector store 5 vector_store = WeaviateVectorStore(weaviate_client = client, index_name="SBCZoning") ----> 6 vector_store.add(nodes) 8 # setting up the storage for the embeddings 9 storage_context = StorageContext.from_defaults(vector_store = vector_store) File ~\anaconda3\lib\site-packages\llama_index\vector_stores\weaviate.py:181, in WeaviateVectorStore.add(self, nodes) 179 with self._client.batch as batch: 180 for node in nodes: ---> 181 add_node( 182 self._client, 183 node, 184 self.index_name, 185 batch=batch, 186 text_key=self.text_key, 187 ) 188 return ids File ~\anaconda3\lib\site-packages\llama_index\vector_stores\weaviate_utils.py:152, in add_node(client, node, class_name, batch, text_key) 149 metadata = {} 150 metadata[text_key] = node.get_content(metadata_mode=MetadataMode.NONE) or "" --> 152 additional_metadata = node_to_metadata_dict( 153 node, remove_text=True, flat_metadata=False 154 ) 155 metadata.update(additional_metadata) 157 vector = node.get_embedding() File ~\anaconda3\lib\site-packages\llama_index\vector_stores\utils.py:46, in node_to_metadata_dict(node, remove_text, text_field, flat_metadata) 43 node_dict["embedding"] = None 45 # dump remainder of node_dict to json string ---> 46 metadata["_node_content"] = json.dumps(node_dict) 48 # store ref doc id at top level to allow metadata filtering 49 # kept for backwards compatibility, will consolidate in future 50 metadata["document_id"] = node.ref_doc_id or "None" # for Chroma File ~\anaconda3\lib\json\__init__.py:231, in dumps(obj, skipkeys, ensure_ascii, check_circular, allow_nan, cls, indent, separators, default, sort_keys, **kw) 226 # cached encoder 227 if (not skipkeys and ensure_ascii and 228 check_circular and allow_nan and 229 cls is None and indent is None and separators is None and 230 default is None and not sort_keys and not kw): --> 231 return _default_encoder.encode(obj) 232 if cls is None: 233 cls = JSONEncoder File ~\anaconda3\lib\json\encoder.py:199, in JSONEncoder.encode(self, o) 195 return encode_basestring(o) 196 # This doesn't pass the iterator directly to ''.join() because the 197 # exceptions aren't as detailed. The list call should be roughly 198 # equivalent to the PySequence_Fast that ''.join() would do. --> 199 chunks = self.iterencode(o, _one_shot=True) 200 if not isinstance(chunks, (list, tuple)): 201 chunks = list(chunks) File ~\anaconda3\lib\json\encoder.py:257, in JSONEncoder.iterencode(self, o, _one_shot) 252 else: 253 _iterencode = _make_iterencode( 254 markers, self.default, _encoder, self.indent, floatstr, 255 self.key_separator, self.item_separator, self.sort_keys, 256 self.skipkeys, _one_shot) --> 257 return _iterencode(o, 0) File ~\anaconda3\lib\json\encoder.py:179, in JSONEncoder.default(self, o) 160 def default(self, o): 161 """Implement this method in a subclass such that it returns 162 a serializable object for ``o``, or calls the base implementation 163 (to raise a ``TypeError``). (...) 177 178 """ --> 179 raise TypeError(f'Object of type {o.__class__.__name__} ' 180 f'is not JSON serializable') TypeError: Object of type WindowsPath is not JSON serializable
注:我是按教程步骤操作的:用PyMuPDFReader加载PDF,SentenceSplitter处理文档并保留源索引,通过MetadataExtractor添加元数据。
报错原因
PyMuPDFReader加载PDF时,会将文件路径以WindowsPath(Path类的子类)对象的形式存入节点元数据。而Weaviate在将节点元数据序列化为JSON时,Python默认的JSON编码器无法处理Path类型对象,只能序列化字符串、数字、列表等基本类型。
解决方法
方法1:加载PDF时直接将路径转为字符串
修改PDF加载代码,把Path对象转为字符串传入:
from llama_index.readers.file import PyMuPDFReader from pathlib import Path # 替换为你的PDF文件路径 file_path = Path("path/to/your/document.pdf") reader = PyMuPDFReader() # 这里将Path对象转为字符串 documents = reader.load_data(file_path=str(file_path))
方法2:批量转换现有节点的Path元数据
如果已经生成了nodes列表,遍历所有节点,把元数据中的Path对象转为字符串:
from pathlib import Path for node in nodes: # 遍历元数据的所有键值对 for key in list(node.metadata.keys()): value = node.metadata[key] if isinstance(value, Path): # 将Path对象转为字符串 node.metadata[key] = str(value)
处理完成后再调用vector_store.add(nodes)即可。
方法3:自定义元数据提取规则(可选)
如果使用MetadataExtractor,可以添加一个自定义处理函数,在提取元数据时自动将路径转为字符串:
from llama_index.node_parser.extractors import MetadataExtractor, BaseExtractor from pathlib import Path class PathToStringExtractor(BaseExtractor): def extract(self, node, **kwargs): metadata = {} for key, value in node.metadata.items(): if isinstance(value, Path): metadata[key] = str(value) return metadata # 初始化MetadataExtractor时添加自定义提取器 metadata_extractor = MetadataExtractor( extractors=[ # 保留你原来的其他提取器 PathToStringExtractor(), ] )
内容的提问来源于stack exchange,提问作者Simon Palmer
相关产品推荐
相关产品推荐

