使用Langchain RecursiveJsonSplitter拆分MongoDB JSON遇索引越界错误求助
问题描述
从MongoDB获取JSON数据后,使用Langchain的RecursiveJsonSplitter时触发IndexError: list index out of range错误,需求是除数据库外不使用任何本地文件。
MongoDB数据检索代码
cursor = collection.find({}) # return json.dumps(cursor, default=json_util.default) json_docs = [json.dumps(doc, default=json_util.default) for doc in cursor] return json_docs
主代码(错误触发点)
# Get data from MongoDB json_data = import_json_files() # Output JSON data to a file with open('data.txt', 'w') as f: json.dump(json_data, f, ensure_ascii=False, indent=4, sort_keys=True, separators=(',', ': ')) with open('client_secrets.json') as f: secrets = json.load(f) os.environ["OPENAI_API_KEY"] = secrets['openai_api_key'] # Convert and ingest documents into vectorstore json_data_str = json.dumps(json_data) documents = json.loads(json_data_str) if not documents: print("No documents found.") exit() text_splitter = RecursiveJsonSplitter(max_chunk_size=1000) json_chunks = text_splitter.split_json(json_data=documents) # 错误触发行 embeddings = OpenAIEmbeddings() db = FAISS.from_documents(json_chunks, embeddings) retriever = db.as_retriever()
错误信息
File "/Users/abcd/Desktop/RAG/simple_example.py", line 38, in <module> json_chunks = text_splitter.split_json(json_data=documents) File "/Library/Frameworks/Python.framework/Versions/3.9/lib/python3.9/site-packages/langchain_text_splitters/json.py", line 89, in split_json chunks = self._json_split(json_data) File "/Library/Frameworks/Python.framework/Versions/3.9/lib/python3.9/site-packages/langchain_text_splitters/json.py", line 76, in _json_split self._set_nested_dict(chunks[-1], current_path, data) File "/Library/Frameworks/Python.framework/Versions/3.9/lib/python3.9/site-packages/langchain_text_splitters/json.py", line 32, in _set_nested_dict d[path[-1]] = value IndexError: list index out of range
问题分析与解决方法
错误原因
你从MongoDB返回的是JSON字符串组成的列表,后续序列化再反序列化后得到的仍是字符串数组。但RecursiveJsonSplitter要求输入的是Python原生的嵌套字典/数组结构,而非字符串数组。当splitter尝试处理字符串时,无法解析出有效的JSON路径,导致内部路径列表为空,触发path[-1]索引越界。
解决方案
1. 修正MongoDB数据获取逻辑
不要将每个文档转为JSON字符串,直接将BSON对象转为Python原生字典:
cursor = collection.find({}) # 使用json_util将BSON转为Python原生字典,而非JSON字符串 json_docs = [json_util._json_convert(doc) for doc in cursor] return json_docs
2. 移除冗余的序列化反序列化步骤
获取到原生Python对象后,无需再执行json.dumps和json.loads:
# Get data from MongoDB documents = import_json_files() if not documents: print("No documents found.") exit() # 移除多余的序列化反序列化代码 # json_data_str = json.dumps(json_data) # documents = json.loads(json_data_str) text_splitter = RecursiveJsonSplitter(max_chunk_size=1000) json_chunks = text_splitter.split_json(json_data=documents) embeddings = OpenAIEmbeddings() db = FAISS.from_documents(json_chunks, embeddings) retriever = db.as_retriever()
3. 移除本地文件写入代码(符合需求)
删除不必要的data.txt写入逻辑,完全避免本地文件依赖:
# 移除以下代码块 # with open('data.txt', 'w') as f: # json.dump(json_data, f, ensure_ascii=False, indent=4, sort_keys=True, separators=(',', ': '))
可行性说明
你的需求完全可行,只要确保传入RecursiveJsonSplitter的是正确的Python原生JSON结构(字典/数组),而非字符串数组,即可正常完成拆分与向量入库流程。
内容的提问来源于stack exchange,提问作者firepower1233
相关产品推荐
相关产品推荐

