如何通过Weaviate实现相似文本匹配并返回关联缓存值?
用Weaviate实现GPT请求缓存的方案
核心思路
不用自定义向量模型,直接复用Weaviate自带的文本向量化模块(比如适配GPT场景的text2vec-openai),把拼接后的GPT请求+响应文本作为向量字段,关联值存在普通属性中。查询时用新请求文本生成向量做相似性检索,通过阈值过滤结果,完全满足松散文本匹配的需求。
具体实现步骤
1. 定义Weaviate类结构
创建一个GPT_Cache类,包含两个核心属性:
cache_key:存储拼接后的GPT请求+响应文本,标记为可向量化字段,由Weaviate自动生成向量cache_value:存储需要缓存的关联值(比如GPT原始响应)
示例Python代码:
import weaviate # 初始化Weaviate客户端 client = weaviate.Client("http://localhost:8080") # 清理已存在的同名类(可选) if client.schema.exists("GPT_Cache"): client.schema.delete_class("GPT_Cache") # 定义类结构 class_config = { "class": "GPT_Cache", "properties": [ { "name": "cache_key", "dataType": ["text"], "vectorizePropertyName": True }, { "name": "cache_value", "dataType": ["text"] } ], "vectorizer": "text2vec-openai" } client.schema.create_class(class_config)
2. 写入缓存数据
调用GPT得到响应后,将请求文本与响应文本拼接为cache_key,和响应值一起存入Weaviate:
def save_to_cache(prompt, gpt_response): # 拼接请求+响应作为缓存键 cache_key_content = f"PROMPT: {prompt} | RESPONSE: {gpt_response}" client.data_object.create( data_object={ "cache_key": cache_key_content, "cache_value": gpt_response }, class_name="GPT_Cache" )
3. 查询相似缓存
新请求到来时,用新prompt发起相似性检索,设置相似度阈值过滤结果,返回匹配的缓存值:
def retrieve_cached_response(new_prompt, similarity_threshold=0.7): # 基于新prompt做相似性查询 query_result = client.query.get( "GPT_Cache", ["cache_value", "_additional { distance }"] ).with_near_text({ "concepts": [new_prompt], "distance": similarity_threshold }).do() # 返回最相似的缓存结果,无匹配则返回None cached_items = query_result["data"]["Get"]["GPT_Cache"] return cached_items[0]["cache_value"] if cached_items else None
4. 调整匹配松紧度
Weaviate的distance参数基于余弦相似度转换:
- 值越小,匹配越严格;值越大,匹配越松散
- 若要更宽松的匹配,可将
similarity_threshold调至0.8甚至0.9,即使拼接文本不完整、不通顺,只要语义相近就能命中
方案优势
- 无需部署自定义BERT模型,直接复用OpenAI Embedding,轻量化且适配GPT语义空间
- 完全满足松散文本匹配需求,无需复杂语言感知逻辑
- 纯Weaviate操作,符合你必须使用该工具的要求
内容的提问来源于stack exchange,提问作者John R Ramsden
相关产品推荐
相关产品推荐

