You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Llama Index结合TimescaleVector存储与加载索引问题求助

解决LlamaIndex + TimescaleVector索引存储与加载问题

错误原因分析

出现KeyError: 'default'的核心原因:

  • 你创建的StorageContext未关联任何VectorStore,是一个空上下文
  • 未正确创建VectorStoreIndex并绑定Timescale向量存储,直接调用persist操作无效
  • ts_vector_store.create_index("aaa")是在Timescale数据库内创建向量查询加速索引,并非保存LlamaIndex的索引配置

正确解决方案

TimescaleVectorStore本身将向量数据存储在Timescale数据库中,无需本地持久化向量。我们只需本地保存LlamaIndex的索引元数据(如嵌入模型配置、索引结构),后续加载时直接连接到Timescale的向量表即可。

步骤1:首次运行 - 创建并保存索引配置(仅需执行一次)

pip install llama_index llama-index-vector-stores-timescalevector llama-index-embeddings-openai

import os
import pandas as pd
from datetime import datetime
from llama_index.core import StorageContext, VectorStoreIndex
from llama_index.core.schema import TextNode
from llama_index.vector_stores.timescalevector import TimescaleVectorStore
from llama_index.embeddings.openai import OpenAIEmbedding

# 设置环境变量
os.environ["OPENAI_API_KEY"] = 'your_openai_api_key'
os.environ["TIMESCALE_SERVICE_URL"] = 'your_timescale_service_url'

# 加载并处理数据
reuters = pd.read_csv('your_file_path')
reuters.columns = ["title", "date", "description"]

# 工具函数(保留原有逻辑)
def create_uuid2(date_string: str):
    if date_string is None:
        return None
    time_format = '%b %d %Y'
    datetime_obj = datetime.strptime(date_string, time_format)
    from timescale_vector import client as timescale_client
    uuid = timescale_client.uuid_from_time(datetime_obj)
    return str(uuid)

def create_date2(input_string: str) -> str:
    if input_string is None:
        return None
    date_object = datetime.strptime(input_string, '%b %d %Y')
    time = "00:00:00"
    timezone_hours = 8
    timezone_minutes = 50
    timestamp_tz_str = f"{date_object.year}-{date_object.month:02}-{date_object.day:02} {time}+{timezone_hours:02}{timezone_minutes:02}"
    return timestamp_tz_str

def create_node2(row):
    record = row.to_dict()
    record_content = f"{record['date']} {record['title']} {record['description']}"
    node = TextNode(
        id_=create_uuid2(str(record["date"])),
        text=record_content,
        metadata={
            "title": record["title"],
            "date": create_date2(str(record["date"])),
        },
    )
    return node

# 创建节点
nodes = [create_node2(row) for _, row in reuters.iterrows()]
embedding_model = OpenAIEmbedding()

# 1. 初始化Timescale向量存储
ts_vector_store = TimescaleVectorStore.from_params(
    service_url=os.environ["TIMESCALE_SERVICE_URL"],
    table_name="reuters_test"
)

# 2. 创建存储上下文,关联Timescale向量存储
storage_context = StorageContext.from_defaults(vector_store=ts_vector_store)

# 3. 创建VectorStoreIndex,自动将节点嵌入并存储到Timescale
index = VectorStoreIndex(
    nodes[:100],
    embedding=embedding_model,
    storage_context=storage_context
)

# 4. 在Timescale内创建向量查询加速索引(可选,提升查询性能)
ts_vector_store.create_index("vector_index")

# 5. 保存索引元数据到本地目录(后续加载用)
index.storage_context.persist(persist_dir="./timescale_index_persist")

步骤2:后续运行 - 加载索引(无需重建)

import os
from llama_index.core import StorageContext, load_index_from_storage
from llama_index.vector_stores.timescalevector import TimescaleVectorStore
from llama_index.embeddings.openai import OpenAIEmbedding

# 设置环境变量
os.environ["OPENAI_API_KEY"] = 'your_openai_api_key'
os.environ["TIMESCALE_SERVICE_URL"] = 'your_timescale_service_url'

# 1. 重新连接到Timescale的向量表
ts_vector_store = TimescaleVectorStore.from_params(
    service_url=os.environ["TIMESCALE_SERVICE_URL"],
    table_name="reuters_test"
)

# 2. 创建存储上下文,关联已有的向量存储与本地元数据
storage_context = StorageContext.from_defaults(
    vector_store=ts_vector_store,
    persist_dir="./timescale_index_persist"
)

# 3. 加载索引
embedding_model = OpenAIEmbedding()
index = load_index_from_storage(
    storage_context,
    embed_model=embedding_model
)

# 正常使用索引查询
query_engine = index.as_query_engine()
response = query_engine.query("你的查询语句")
print(response)

关键注意事项

  • 向量数据存储位置:所有向量和节点元数据都保存在Timescale数据库中,本地目录仅存储索引配置信息
  • 避免重复导入:首次运行后,后续加载索引时不要再次调用ts_vector_store.add(nodes),否则会重复插入数据
  • Timescale内的索引:ts_vector_store.create_index()是为数据库创建查询加速索引,建议仅执行一次
  • 版本一致性:确保每次运行的依赖包版本与首次创建索引时一致

内容的提问来源于stack exchange,提问作者Gianluca Baglini

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 12:10:05