无法用Llama Index加载S3索引文件,报错'str'无get_doc_id属性
问题描述
尝试加载保存在S3的Llama索引JSON文件时,触发错误:
Error: 'str' object has no attribute 'get_doc_id'
当前使用的代码是直接从S3读取JSON内容,再通过VectorStoreIndex.from_documents()加载,代码如下:
import boto3 import json from llama_index import VectorStoreIndex s3 = boto3.client('s3') def read_index_file(bucket_name, file_name): try: response = s3.get_object(Bucket=bucket_name, Key=file_name) content = response['Body'].read().decode('utf-8') index_data = json.loads(content) return index_data except Exception as e: print("Error reading index file:", str(e)) return None bucket_name = 'your_bucket_name' file_name = 'your_file_name.json' index_data = read_index_file(bucket_name, file_name) index = VectorStoreIndex.from_documents(index_data)
确认S3内容可正常读取,但加载索引对象失败,旧版本llama_index可正常运行该逻辑。
问题原因
VectorStoreIndex.from_documents()要求传入Document对象的列表,而非JSON解析后的字典或字符串数据。新版本Llama Index改用标准化的StorageContext机制处理索引的持久化与加载,直接读取JSON解析后的数据无法被识别为合法输入。
解决方法
改用Llama Index官方的存储加载API,结合S3文件系统配置存储上下文来加载索引:
步骤1:确保保存索引时使用官方持久化方法
如果是自行保存的索引,需用StorageContext的persist()方法指定S3路径,示例代码:
from llama_index import VectorStoreIndex, StorageContext from llama_index.vector_stores import SimpleVectorStore from llama_index.storage.docstore import SimpleDocumentStore from llama_index.storage.index_store import SimpleIndexStore from llama_index.fs import S3FileSystem # 假设已构建好index s3_fs = S3FileSystem(bucket_name="your_bucket_name") storage_context = StorageContext.from_defaults( docstore=SimpleDocumentStore.from_persist_dir(persist_dir="s3://your_bucket_name/index_dir", fs=s3_fs), vector_store=SimpleVectorStore.from_persist_dir(persist_dir="s3://your_bucket_name/index_dir", fs=s3_fs), index_store=SimpleIndexStore.from_persist_dir(persist_dir="s3://your_bucket_name/index_dir", fs=s3_fs), ) # 保存索引到S3 index.storage_context.persist(persist_dir="s3://your_bucket_name/index_dir", fs=s3_fs)
步骤2:加载S3中的索引
直接通过StorageContext从S3路径加载:
from llama_index import load_index_from_storage from llama_index.storage.storage_context import StorageContext from llama_index.fs import S3FileSystem s3_fs = S3FileSystem(bucket_name="your_bucket_name") # 从S3的索引目录加载存储上下文 storage_context = StorageContext.from_defaults( persist_dir="s3://your_bucket_name/index_dir", fs=s3_fs ) # 加载索引对象 index = load_index_from_storage(storage_context)
补充说明
旧版本llama_index可能允许直接序列化/反序列化JSON加载索引,但新版本已废弃该方式,统一使用StorageContext和load_index_from_storage()处理索引的持久化与加载,保障兼容性与扩展性。
内容的提问来源于stack exchange,提问作者beachCode
相关产品推荐
相关产品推荐

