使用LangChain+Weaviate调用RetrievalQAWithSourcesChain遇ValueError求助
问题:LangChain结合Weaviate使用RetrievalQAWithSourcesChain触发source元数据缺失错误
使用LangChain框架、向量数据库(Weaviate/FAISS)搭配RetrievalQAWithSourcesChain链时,FAISS可正常返回结果,但切换到Weaviate时触发报错:
ValueError: Document prompt requires documents to have metadata variables: ['source']. Received document with missing metadata: ['source']
已在Weaviate的Products类中定义source属性并插入对应数据,调整prompt后问题仍未解决。
主业务代码
from langchain.vectorstores.weaviate import Weaviate from langchain.llms import OpenAI from langchain.chains import RetrievalQAWithSourcesChain import weaviate from langchain.prompts.prompt import PromptTemplate from langchain.vectorstores import FAISS from langchain.embeddings import OpenAIEmbeddings # API Key needs to be passed in playground OPEN_API_KEY="sk-xxxxx" client = weaviate.Client( url="https://xxxxx.weaviate.network", additional_headers={ "X-OpenAI-Api-Key": OPEN_API_KEY } ) vectorstore = Weaviate(client, "Products", "description") # vectorstore = FAISS.load_local( # "./working_fas", # OpenAIEmbeddings(openai_api_key=OPEN_API_KEY) # ) llm = OpenAI(model_name="text-davinci-003", temperature=0, max_tokens=200, openai_api_key=OPEN_API_KEY) template = """ Return product and price information -------------------- {summaries} """ prompt = PromptTemplate( input_variables=["summaries"], template=template, ) chain = RetrievalQAWithSourcesChain.from_chain_type(llm=llm, retriever=vectorstore.as_retriever(), return_source_documents=False, chain_type_kwargs = {"prompt": prompt} ) result = chain("suggest me an watch", return_only_outputs=True) print(result)
Weaviate类定义代码
# Define class and property definitions for products class_def = { "class": "Products", "description": "Products", "properties": [ { "dataType": ["text"], "description": "product category", "name": "category" }, { "name": "sku", "description": "product sku", "dataType": ["text"] }, { "dataType": ["text"], "name": "product", "description": "product name" }, { "dataType": ["text"], "name": "description", "description": "product description" }, { "name": "price", "dataType": ["number"], "description": "product price" }, { "name": "breadcrumb", "dataType": ["text"], "description": "product breadcrumb" }, { "name": "source", "dataType": ["text"], "description": "product url", }, { "name": "money_back", "dataType": ["boolean"], "description": "money_back / refund available for the product" }, { "name": "rating", "dataType": ["number"], "description": "product rating" }, { "name": "total_reviews", "dataType": ["int"], "description": "product total_reviews" }, { "name": "tags", "dataType": ["text"], "description": "product tags" }, { "name": "type", "dataType": ["text"], "description": "product type" } ], "vectorizer": "text2vec-openai", }
Weaviate类创建与数据插入代码
# Create Class client.schema.create_class(class_def)
# Insert datas into class import pandas as pd import time df = pd.read_csv("testing.csv") print(len(df)) for index,row in df.iterrows(): time.sleep(1) properties = { "category": row["category"], "sku": row["sku"], "product": row["product"], "description": row["description"], "price": row["price"], "breadcrumb": row["breadcrumb"], "source": row["source"], "money_back": row["money_back"], "rating": row["rating"], "total_reviews": row["total_reviews"], "tags": row["tags"], "type": row["type"], } print(properties) client.data_object.create(properties, "Products") time.sleep(1)
解决方案
问题核心是LangChain的Weaviate向量存储实现默认不会自动将Weaviate对象的source属性映射到Document的metadata中,需显式指定要提取的元数据字段:
- 修改向量存储初始化代码
添加attributes参数,指定需要同步为Document metadata的字段,至少包含source:
vectorstore = Weaviate(client, "Products", "description", attributes=["source", "product", "price"])
- 验证Weaviate中source字段数据
执行以下代码检查数据是否正确插入:
# 查询一条Products数据,验证source字段 result = client.query.get("Products", ["source", "product"]).do() print(result)
若返回结果中source为空,需检查CSV文件的source列是否包含有效数据。
内容的提问来源于stack exchange,提问作者Syed Mujeeb H
相关产品推荐
相关产品推荐

