You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Langchain RecursiveJsonSplitter拆分MongoDB JSON遇索引越界错误求助

问题描述

从MongoDB获取JSON数据后,使用Langchain的RecursiveJsonSplitter时触发IndexError: list index out of range错误,需求是除数据库外不使用任何本地文件。

MongoDB数据检索代码

cursor = collection.find({})

# return json.dumps(cursor, default=json_util.default)

json_docs = [json.dumps(doc, default=json_util.default) for doc in cursor]
return json_docs

主代码(错误触发点)

# Get data from MongoDB
json_data = import_json_files()

# Output JSON data to a file
with open('data.txt', 'w') as f:
  json.dump(json_data, f, ensure_ascii=False, indent=4, sort_keys=True, separators=(',', ': '))

with open('client_secrets.json') as f:
    secrets = json.load(f)

os.environ["OPENAI_API_KEY"] = secrets['openai_api_key']


# Convert and ingest documents into vectorstore
json_data_str = json.dumps(json_data)
documents = json.loads(json_data_str)

if not documents:
    print("No documents found.")
    exit()

text_splitter = RecursiveJsonSplitter(max_chunk_size=1000)
json_chunks = text_splitter.split_json(json_data=documents)  # 错误触发行
embeddings = OpenAIEmbeddings()
db = FAISS.from_documents(json_chunks, embeddings)

retriever = db.as_retriever()

错误信息

File "/Users/abcd/Desktop/RAG/simple_example.py", line 38, in <module>
    json_chunks = text_splitter.split_json(json_data=documents)
  File "/Library/Frameworks/Python.framework/Versions/3.9/lib/python3.9/site-packages/langchain_text_splitters/json.py", line 89, in split_json
    chunks = self._json_split(json_data)
  File "/Library/Frameworks/Python.framework/Versions/3.9/lib/python3.9/site-packages/langchain_text_splitters/json.py", line 76, in _json_split
    self._set_nested_dict(chunks[-1], current_path, data)
  File "/Library/Frameworks/Python.framework/Versions/3.9/lib/python3.9/site-packages/langchain_text_splitters/json.py", line 32, in _set_nested_dict
    d[path[-1]] = value
IndexError: list index out of range
问题分析与解决方法

错误原因

你从MongoDB返回的是JSON字符串组成的列表,后续序列化再反序列化后得到的仍是字符串数组。但RecursiveJsonSplitter要求输入的是Python原生的嵌套字典/数组结构,而非字符串数组。当splitter尝试处理字符串时,无法解析出有效的JSON路径,导致内部路径列表为空,触发path[-1]索引越界。

解决方案

1. 修正MongoDB数据获取逻辑

不要将每个文档转为JSON字符串,直接将BSON对象转为Python原生字典:

cursor = collection.find({})
# 使用json_util将BSON转为Python原生字典,而非JSON字符串
json_docs = [json_util._json_convert(doc) for doc in cursor]
return json_docs

2. 移除冗余的序列化反序列化步骤

获取到原生Python对象后,无需再执行json.dumps和json.loads:

# Get data from MongoDB
documents = import_json_files()

if not documents:
    print("No documents found.")
    exit()

# 移除多余的序列化反序列化代码
# json_data_str = json.dumps(json_data)
# documents = json.loads(json_data_str)

text_splitter = RecursiveJsonSplitter(max_chunk_size=1000)
json_chunks = text_splitter.split_json(json_data=documents)
embeddings = OpenAIEmbeddings()
db = FAISS.from_documents(json_chunks, embeddings)

retriever = db.as_retriever()

3. 移除本地文件写入代码(符合需求)

删除不必要的data.txt写入逻辑,完全避免本地文件依赖:

# 移除以下代码块
# with open('data.txt', 'w') as f:
#   json.dump(json_data, f, ensure_ascii=False, indent=4, sort_keys=True, separators=(',', ': '))

可行性说明

你的需求完全可行,只要确保传入RecursiveJsonSplitter的是正确的Python原生JSON结构(字典/数组),而非字符串数组,即可正常完成拆分与向量入库流程。

内容的提问来源于stack exchange,提问作者firepower1233

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 13:51:16