You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法用Llama Index加载S3索引文件,报错'str'无get_doc_id属性

问题描述

尝试加载保存在S3的Llama索引JSON文件时,触发错误:

Error: 'str' object has no attribute 'get_doc_id'

当前使用的代码是直接从S3读取JSON内容,再通过VectorStoreIndex.from_documents()加载,代码如下:

import boto3
import json
from llama_index import VectorStoreIndex

s3 = boto3.client('s3')

def read_index_file(bucket_name, file_name):
    try:
        response = s3.get_object(Bucket=bucket_name, Key=file_name)
        content = response['Body'].read().decode('utf-8')
        index_data = json.loads(content)
        return index_data
    except Exception as e:
        print("Error reading index file:", str(e))
        return None

bucket_name = 'your_bucket_name'
file_name = 'your_file_name.json'
index_data = read_index_file(bucket_name, file_name)

index = VectorStoreIndex.from_documents(index_data)

确认S3内容可正常读取,但加载索引对象失败,旧版本llama_index可正常运行该逻辑。

问题原因

VectorStoreIndex.from_documents()要求传入Document对象的列表,而非JSON解析后的字典或字符串数据。新版本Llama Index改用标准化的StorageContext机制处理索引的持久化与加载,直接读取JSON解析后的数据无法被识别为合法输入。

解决方法

改用Llama Index官方的存储加载API,结合S3文件系统配置存储上下文来加载索引:

步骤1:确保保存索引时使用官方持久化方法

如果是自行保存的索引,需用StorageContext的persist()方法指定S3路径,示例代码:

from llama_index import VectorStoreIndex, StorageContext
from llama_index.vector_stores import SimpleVectorStore
from llama_index.storage.docstore import SimpleDocumentStore
from llama_index.storage.index_store import SimpleIndexStore
from llama_index.fs import S3FileSystem

# 假设已构建好index
s3_fs = S3FileSystem(bucket_name="your_bucket_name")
storage_context = StorageContext.from_defaults(
    docstore=SimpleDocumentStore.from_persist_dir(persist_dir="s3://your_bucket_name/index_dir", fs=s3_fs),
    vector_store=SimpleVectorStore.from_persist_dir(persist_dir="s3://your_bucket_name/index_dir", fs=s3_fs),
    index_store=SimpleIndexStore.from_persist_dir(persist_dir="s3://your_bucket_name/index_dir", fs=s3_fs),
)
# 保存索引到S3
index.storage_context.persist(persist_dir="s3://your_bucket_name/index_dir", fs=s3_fs)

步骤2:加载S3中的索引

直接通过StorageContext从S3路径加载:

from llama_index import load_index_from_storage
from llama_index.storage.storage_context import StorageContext
from llama_index.fs import S3FileSystem

s3_fs = S3FileSystem(bucket_name="your_bucket_name")
# 从S3的索引目录加载存储上下文
storage_context = StorageContext.from_defaults(
    persist_dir="s3://your_bucket_name/index_dir",
    fs=s3_fs
)
# 加载索引对象
index = load_index_from_storage(storage_context)

补充说明

旧版本llama_index可能允许直接序列化/反序列化JSON加载索引,但新版本已废弃该方式,统一使用StorageContext和load_index_from_storage()处理索引的持久化与加载,保障兼容性与扩展性。

内容的提问来源于stack exchange,提问作者beachCode

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 02:52:27