Llama Index结合TimescaleVector存储与加载索引问题求助
解决LlamaIndex + TimescaleVector索引存储与加载问题
错误原因分析
出现KeyError: 'default'的核心原因:
- 你创建的
StorageContext未关联任何VectorStore,是一个空上下文 - 未正确创建
VectorStoreIndex并绑定Timescale向量存储,直接调用persist操作无效 ts_vector_store.create_index("aaa")是在Timescale数据库内创建向量查询加速索引,并非保存LlamaIndex的索引配置
正确解决方案
TimescaleVectorStore本身将向量数据存储在Timescale数据库中,无需本地持久化向量。我们只需本地保存LlamaIndex的索引元数据(如嵌入模型配置、索引结构),后续加载时直接连接到Timescale的向量表即可。
步骤1:首次运行 - 创建并保存索引配置(仅需执行一次)
pip install llama_index llama-index-vector-stores-timescalevector llama-index-embeddings-openai import os import pandas as pd from datetime import datetime from llama_index.core import StorageContext, VectorStoreIndex from llama_index.core.schema import TextNode from llama_index.vector_stores.timescalevector import TimescaleVectorStore from llama_index.embeddings.openai import OpenAIEmbedding # 设置环境变量 os.environ["OPENAI_API_KEY"] = 'your_openai_api_key' os.environ["TIMESCALE_SERVICE_URL"] = 'your_timescale_service_url' # 加载并处理数据 reuters = pd.read_csv('your_file_path') reuters.columns = ["title", "date", "description"] # 工具函数(保留原有逻辑) def create_uuid2(date_string: str): if date_string is None: return None time_format = '%b %d %Y' datetime_obj = datetime.strptime(date_string, time_format) from timescale_vector import client as timescale_client uuid = timescale_client.uuid_from_time(datetime_obj) return str(uuid) def create_date2(input_string: str) -> str: if input_string is None: return None date_object = datetime.strptime(input_string, '%b %d %Y') time = "00:00:00" timezone_hours = 8 timezone_minutes = 50 timestamp_tz_str = f"{date_object.year}-{date_object.month:02}-{date_object.day:02} {time}+{timezone_hours:02}{timezone_minutes:02}" return timestamp_tz_str def create_node2(row): record = row.to_dict() record_content = f"{record['date']} {record['title']} {record['description']}" node = TextNode( id_=create_uuid2(str(record["date"])), text=record_content, metadata={ "title": record["title"], "date": create_date2(str(record["date"])), }, ) return node # 创建节点 nodes = [create_node2(row) for _, row in reuters.iterrows()] embedding_model = OpenAIEmbedding() # 1. 初始化Timescale向量存储 ts_vector_store = TimescaleVectorStore.from_params( service_url=os.environ["TIMESCALE_SERVICE_URL"], table_name="reuters_test" ) # 2. 创建存储上下文,关联Timescale向量存储 storage_context = StorageContext.from_defaults(vector_store=ts_vector_store) # 3. 创建VectorStoreIndex,自动将节点嵌入并存储到Timescale index = VectorStoreIndex( nodes[:100], embedding=embedding_model, storage_context=storage_context ) # 4. 在Timescale内创建向量查询加速索引(可选,提升查询性能) ts_vector_store.create_index("vector_index") # 5. 保存索引元数据到本地目录(后续加载用) index.storage_context.persist(persist_dir="./timescale_index_persist")
步骤2:后续运行 - 加载索引(无需重建)
import os from llama_index.core import StorageContext, load_index_from_storage from llama_index.vector_stores.timescalevector import TimescaleVectorStore from llama_index.embeddings.openai import OpenAIEmbedding # 设置环境变量 os.environ["OPENAI_API_KEY"] = 'your_openai_api_key' os.environ["TIMESCALE_SERVICE_URL"] = 'your_timescale_service_url' # 1. 重新连接到Timescale的向量表 ts_vector_store = TimescaleVectorStore.from_params( service_url=os.environ["TIMESCALE_SERVICE_URL"], table_name="reuters_test" ) # 2. 创建存储上下文,关联已有的向量存储与本地元数据 storage_context = StorageContext.from_defaults( vector_store=ts_vector_store, persist_dir="./timescale_index_persist" ) # 3. 加载索引 embedding_model = OpenAIEmbedding() index = load_index_from_storage( storage_context, embed_model=embedding_model ) # 正常使用索引查询 query_engine = index.as_query_engine() response = query_engine.query("你的查询语句") print(response)
关键注意事项
- 向量数据存储位置:所有向量和节点元数据都保存在Timescale数据库中,本地目录仅存储索引配置信息
- 避免重复导入:首次运行后,后续加载索引时不要再次调用
ts_vector_store.add(nodes),否则会重复插入数据 - Timescale内的索引:
ts_vector_store.create_index()是为数据库创建查询加速索引,建议仅执行一次 - 版本一致性:确保每次运行的依赖包版本与首次创建索引时一致
内容的提问来源于stack exchange,提问作者Gianluca Baglini
相关产品推荐
相关产品推荐

