LangChain聊天机器人:替换ChromaDB为Weaviate遇报错及架构优化咨询
问题解答
一、是否可以跳过Azure Cognitive Search直接对Azure SQL表建立索引?
完全可以,这种方式能简化流程、减少中间环节的性能损耗,推荐采用。具体有两种可行方案:
- 方案1:通过LangChain直接读取Azure SQL数据并导入Weaviate
用LangChain的SQLDatabase连接Azure SQL,查询profiles表的所有数据,将每行数据拼接成符合要求的content字段,转换为LangChain的Document对象后直接导入Weaviate。示例代码如下:from langchain.sql_database import SQLDatabase from langchain.docstore.document import Document # 配置Azure SQL连接字符串 db_uri = "mssql+pyodbc://<用户名>:<密码>@<服务器名>.database.windows.net:1433/<数据库名>?driver=ODBC+Driver+17+for+SQL+Server" db = SQLDatabase.from_uri(db_uri) # 查询profiles表所有数据 query_result = db.run("SELECT * FROM profiles") # 转换为LangChain Document对象 docs = [] for row in query_result: # 根据实际字段拼接content,保持和之前一致的格式 content = f"姓名:{row['name']},年龄:{row['age']},爱好:{row['hobbies']}" docs.append(Document(page_content=content, metadata={"id": row["id"]})) # 导入Weaviate作为向量存储 vectorstore = Weaviate.from_documents(docs, embeddings, weaviate_url=WEAVIATE_URL) - 方案2:使用Weaviate原生Azure SQL连接器
Weaviate支持直接对接Azure SQL数据源,通过配置Weaviate的数据源连接器和向量生成器,可自动同步Azure SQL中的数据并构建向量索引,无需LangChain中转,性能表现更优。
二、当前Weaviate报错的解决方法
报错原因是从Azure Cognitive Search返回的文档元数据中包含@search.score字段,而Weaviate的属性名不允许包含@符号。只需在导入Weaviate前移除该字段即可:
docs = get_relevant_documents(retriever, user_input) # 移除非法的@search.score字段 for doc in docs: doc.metadata.pop("@search.score", None) # 导入Weaviate vectorstore = from_documentsWeaviate(docs, embeddings)
内容的提问来源于stack exchange,提问作者Luis Valencia
相关产品推荐
相关产品推荐

