基于FAISS与Ollama的RAG如何实现商品性别维度严格过滤?
解决RAG系统中商品严格过滤的方案
1. 给商品添加结构化元数据
首先必须为商品目录补充明确的结构化属性,比如gender(female/male/unisex)、color、category等,示例元数据格式:
{ "product_id": "001", "description": "Women's brown wool scarf - warm winter accessory", "gender": "female", "color": "brown", "category": "scarf" }
生成嵌入时,将商品描述与元数据拼接作为输入(如description + f" Gender: {gender}, Color: {color}, Category: {category}"),让嵌入向量包含结构化属性信息。
2. 实现硬过滤逻辑
向量检索的语义匹配可能返回相似但不符合属性的商品,必须通过硬过滤确保结果严格匹配查询条件,以下是两种适配FAISS的实现方式:
方式一:基于LangChain的FAISS元数据过滤
如果使用LangChain封装FAISS,可直接在检索时传入过滤参数:
from langchain.embeddings import HuggingFaceEmbeddings from langchain.vectorstores import FAISS from langchain.docstore.document import Document # 初始化嵌入模型 embeddings = HuggingFaceEmbeddings(model_name="all-mpnet-base-v2") # 构建带元数据的商品文档 docs = [ Document( page_content="Women's brown wool scarf - soft and warm for winter", metadata={"gender": "female", "color": "brown", "category": "scarf"} ), Document( page_content="Men's brown cotton scarf - durable for outdoor use", metadata={"gender": "male", "color": "brown", "category": "scarf"} ) ] # 创建FAISS索引 db = FAISS.from_documents(docs, embeddings) # 解析查询得到过滤条件(可通过Ollama提取或规则匹配) filter_dict = {"gender": "female", "color": "brown", "category": "scarf"} # 带过滤的检索 retrieved_docs = db.similarity_search( "Looking for a brown scarf for women", k=3, filter=filter_dict ) # 无匹配时返回空结果 if not retrieved_docs: print("No matching products found.") else: # 将结果传入Ollama生成回答 pass
方式二:自定义元数据索引+检索后过滤
若不使用LangChain,可自行维护元数据与FAISS索引的映射,检索后过滤不符合条件的结果:
import faiss import numpy as np from sentence_transformers import SentenceTransformer # 初始化嵌入模型 model = SentenceTransformer("all-mpnet-base-v2") # 商品数据与元数据 products = [ {"id": 0, "desc": "Women's brown wool scarf", "gender": "female", "color": "brown", "category": "scarf"}, {"id": 1, "desc": "Men's brown cotton scarf", "gender": "male", "color": "brown", "category": "scarf"} ] # 生成嵌入向量 texts = [p["desc"] + f" Gender: {p['gender']}, Color: {p['color']}, Category: {p['category']}" for p in products] embeddings = model.encode(texts) # 构建FAISS索引 index = faiss.IndexFlatL2(embeddings.shape[1]) index.add(np.array(embeddings)) # 处理查询 query = "Looking for a brown scarf for women" query_embedding = model.encode([query]) # 检索Top-K结果 k = 3 distances, indices = index.search(query_embedding, k) # 过滤符合条件的商品 filtered_products = [] for idx in indices[0]: product = products[idx] if product["gender"] == "female" and product["color"] == "brown" and product["category"] == "scarf": filtered_products.append(product) # 输出结果 if not filtered_products: print("No matching products found.") else: print(filtered_products)
3. 优化查询条件解析
用Ollama实现自然语言查询的属性提取,确保过滤条件准确:
from ollama import generate prompt = """ Extract product attributes from the user query below, output in JSON format with keys: gender, color, category. Omit keys if the attribute is not mentioned. User query: "{query}" Example output: {"gender": "female", "color": "brown", "category": "scarf"} """ # 提取过滤条件 response = generate(model="llama3", prompt=prompt.format(query="Looking for a brown scarf for women")) # 解析response得到filter_dict
4. 检索后二次校验
为确保结果完全匹配,可让Ollama对检索到的商品做二次校验:
# 假设已得到retrieved_docs check_prompt = """ Check if the product matches the user query. Reply ONLY with "YES" or "NO". User query: "{query}" Product description: {product_desc} """ # 过滤不符合的商品 valid_docs = [] for doc in retrieved_docs: check_result = generate(model="llama3", prompt=check_prompt.format(query=query, product_desc=doc.page_content)) if check_result.strip() == "YES": valid_docs.append(doc) if not valid_docs: print("No matching products found.")
内容的提问来源于stack exchange,提问作者Advait Shendage
相关产品推荐
相关产品推荐

