You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用llama-index向Weaviate加载Nodes时遇WindowsPath序列化错误

解决Weaviate向量存储导入时的WindowsPath JSON序列化错误

问题场景

我在跟随llama-index的《Building Data Ingestion From Scratch》教程实现数据导入,原教程使用Pinecone的代码如下:

from llama_index.vector_stores import PineconeVectorStore
vector_store = PineconeVectorStore(pinecone_index=pinecone_index)
vector_store.add(nodes)

我将其替换为Weaviate的实现:

from llama_index.vector_stores import WeaviateVectorStore
# 构建向量存储
vector_store = WeaviateVectorStore(weaviate_client = client, index_name="SBCZoning")
vector_store.add(nodes)

运行时触发TypeError,提示WindowsPath类型对象无法JSON序列化,完整报错堆栈:

TypeError                                 Traceback (most recent call last)
Input In [73], in <cell line: 6>()
      1 # nodes_to_parse = SimpleNodeParser.get_nodes_from_documents(nodes)
      2 # nodes_to_parse = parser.get_nodes_from_documents(nodes)
      3 
      4 # construct vector store
      5 vector_store = WeaviateVectorStore(weaviate_client = client, index_name="SBCZoning")
----> 6 vector_store.add(nodes)
      8 # setting up the storage for the embeddings
      9 storage_context = StorageContext.from_defaults(vector_store = vector_store)

File ~\anaconda3\lib\site-packages\llama_index\vector_stores\weaviate.py:181, in WeaviateVectorStore.add(self, nodes)
    179 with self._client.batch as batch:
    180     for node in nodes:
---> 181         add_node(
    182             self._client,
    183             node,
    184             self.index_name,
    185             batch=batch,
    186             text_key=self.text_key,
    187         )
    188 return ids

File ~\anaconda3\lib\site-packages\llama_index\vector_stores\weaviate_utils.py:152, in add_node(client, node, class_name, batch, text_key)
    149 metadata = {}
    150 metadata[text_key] = node.get_content(metadata_mode=MetadataMode.NONE) or ""
--> 152 additional_metadata = node_to_metadata_dict(
    153     node, remove_text=True, flat_metadata=False
    154 )
    155 metadata.update(additional_metadata)
    157 vector = node.get_embedding()

File ~\anaconda3\lib\site-packages\llama_index\vector_stores\utils.py:46, in node_to_metadata_dict(node, remove_text, text_field, flat_metadata)
     43 node_dict["embedding"] = None
     45 # dump remainder of node_dict to json string
---> 46 metadata["_node_content"] = json.dumps(node_dict)
     48 # store ref doc id at top level to allow metadata filtering
     49 # kept for backwards compatibility, will consolidate in future
     50 metadata["document_id"] = node.ref_doc_id or "None"  # for Chroma

File ~\anaconda3\lib\json\__init__.py:231, in dumps(obj, skipkeys, ensure_ascii, check_circular, allow_nan, cls, indent, separators, default, sort_keys, **kw)
    226 # cached encoder
    227 if (not skipkeys and ensure_ascii and
    228     check_circular and allow_nan and
    229     cls is None and indent is None and separators is None and
    230     default is None and not sort_keys and not kw):
--> 231     return _default_encoder.encode(obj)
    232 if cls is None:
    233     cls = JSONEncoder

File ~\anaconda3\lib\json\encoder.py:199, in JSONEncoder.encode(self, o)
    195         return encode_basestring(o)
    196 # This doesn't pass the iterator directly to ''.join() because the
    197 # exceptions aren't as detailed.  The list call should be roughly
    198 # equivalent to the PySequence_Fast that ''.join() would do.
--> 199 chunks = self.iterencode(o, _one_shot=True)
    200 if not isinstance(chunks, (list, tuple)):
    201     chunks = list(chunks)

File ~\anaconda3\lib\json\encoder.py:257, in JSONEncoder.iterencode(self, o, _one_shot)
    252 else:
    253     _iterencode = _make_iterencode(
    254         markers, self.default, _encoder, self.indent, floatstr,
    255         self.key_separator, self.item_separator, self.sort_keys,
    256         self.skipkeys, _one_shot)
--> 257 return _iterencode(o, 0)

File ~\anaconda3\lib\json\encoder.py:179, in JSONEncoder.default(self, o)
    160 def default(self, o):
    161     """Implement this method in a subclass such that it returns
    162     a serializable object for ``o``, or calls the base implementation
    163     (to raise a ``TypeError``).
   (...)
    177 
    178     """
--> 179     raise TypeError(f'Object of type {o.__class__.__name__} '
    180                     f'is not JSON serializable')

TypeError: Object of type WindowsPath is not JSON serializable

注:我是按教程步骤操作的:用PyMuPDFReader加载PDF,SentenceSplitter处理文档并保留源索引,通过MetadataExtractor添加元数据。

报错原因

PyMuPDFReader加载PDF时,会将文件路径以WindowsPath(Path类的子类)对象的形式存入节点元数据。而Weaviate在将节点元数据序列化为JSON时,Python默认的JSON编码器无法处理Path类型对象,只能序列化字符串、数字、列表等基本类型。

解决方法

方法1:加载PDF时直接将路径转为字符串

修改PDF加载代码,把Path对象转为字符串传入:

from llama_index.readers.file import PyMuPDFReader
from pathlib import Path

# 替换为你的PDF文件路径
file_path = Path("path/to/your/document.pdf")
reader = PyMuPDFReader()
# 这里将Path对象转为字符串
documents = reader.load_data(file_path=str(file_path))

方法2:批量转换现有节点的Path元数据

如果已经生成了nodes列表,遍历所有节点,把元数据中的Path对象转为字符串:

from pathlib import Path

for node in nodes:
    # 遍历元数据的所有键值对
    for key in list(node.metadata.keys()):
        value = node.metadata[key]
        if isinstance(value, Path):
            # 将Path对象转为字符串
            node.metadata[key] = str(value)

处理完成后再调用vector_store.add(nodes)即可。

方法3:自定义元数据提取规则(可选)

如果使用MetadataExtractor,可以添加一个自定义处理函数,在提取元数据时自动将路径转为字符串:

from llama_index.node_parser.extractors import MetadataExtractor, BaseExtractor
from pathlib import Path

class PathToStringExtractor(BaseExtractor):
    def extract(self, node, **kwargs):
        metadata = {}
        for key, value in node.metadata.items():
            if isinstance(value, Path):
                metadata[key] = str(value)
        return metadata

# 初始化MetadataExtractor时添加自定义提取器
metadata_extractor = MetadataExtractor(
    extractors=[
        # 保留你原来的其他提取器
        PathToStringExtractor(),
    ]
)

内容的提问来源于stack exchange,提问作者Simon Palmer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 11:04:56