You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于FAISS与Ollama的RAG如何实现商品性别维度严格过滤?

解决RAG系统中商品严格过滤的方案

1. 给商品添加结构化元数据

首先必须为商品目录补充明确的结构化属性,比如gender(female/male/unisex)、color、category等,示例元数据格式:

{
    "product_id": "001",
    "description": "Women's brown wool scarf - warm winter accessory",
    "gender": "female",
    "color": "brown",
    "category": "scarf"
}

生成嵌入时,将商品描述与元数据拼接作为输入(如description + f" Gender: {gender}, Color: {color}, Category: {category}"),让嵌入向量包含结构化属性信息。

2. 实现硬过滤逻辑

向量检索的语义匹配可能返回相似但不符合属性的商品,必须通过硬过滤确保结果严格匹配查询条件,以下是两种适配FAISS的实现方式:

方式一:基于LangChain的FAISS元数据过滤

如果使用LangChain封装FAISS,可直接在检索时传入过滤参数:

from langchain.embeddings import HuggingFaceEmbeddings
from langchain.vectorstores import FAISS
from langchain.docstore.document import Document

# 初始化嵌入模型
embeddings = HuggingFaceEmbeddings(model_name="all-mpnet-base-v2")

# 构建带元数据的商品文档
docs = [
    Document(
        page_content="Women's brown wool scarf - soft and warm for winter",
        metadata={"gender": "female", "color": "brown", "category": "scarf"}
    ),
    Document(
        page_content="Men's brown cotton scarf - durable for outdoor use",
        metadata={"gender": "male", "color": "brown", "category": "scarf"}
    )
]

# 创建FAISS索引
db = FAISS.from_documents(docs, embeddings)

# 解析查询得到过滤条件(可通过Ollama提取或规则匹配)
filter_dict = {"gender": "female", "color": "brown", "category": "scarf"}

# 带过滤的检索
retrieved_docs = db.similarity_search(
    "Looking for a brown scarf for women",
    k=3,
    filter=filter_dict
)

# 无匹配时返回空结果
if not retrieved_docs:
    print("No matching products found.")
else:
    # 将结果传入Ollama生成回答
    pass

方式二:自定义元数据索引+检索后过滤

若不使用LangChain,可自行维护元数据与FAISS索引的映射,检索后过滤不符合条件的结果:

import faiss
import numpy as np
from sentence_transformers import SentenceTransformer

# 初始化嵌入模型
model = SentenceTransformer("all-mpnet-base-v2")

# 商品数据与元数据
products = [
    {"id": 0, "desc": "Women's brown wool scarf", "gender": "female", "color": "brown", "category": "scarf"},
    {"id": 1, "desc": "Men's brown cotton scarf", "gender": "male", "color": "brown", "category": "scarf"}
]

# 生成嵌入向量
texts = [p["desc"] + f" Gender: {p['gender']}, Color: {p['color']}, Category: {p['category']}" for p in products]
embeddings = model.encode(texts)

# 构建FAISS索引
index = faiss.IndexFlatL2(embeddings.shape[1])
index.add(np.array(embeddings))

# 处理查询
query = "Looking for a brown scarf for women"
query_embedding = model.encode([query])

# 检索Top-K结果
k = 3
distances, indices = index.search(query_embedding, k)

# 过滤符合条件的商品
filtered_products = []
for idx in indices[0]:
    product = products[idx]
    if product["gender"] == "female" and product["color"] == "brown" and product["category"] == "scarf":
        filtered_products.append(product)

# 输出结果
if not filtered_products:
    print("No matching products found.")
else:
    print(filtered_products)

3. 优化查询条件解析

用Ollama实现自然语言查询的属性提取,确保过滤条件准确:

from ollama import generate

prompt = """
Extract product attributes from the user query below, output in JSON format with keys: gender, color, category.
Omit keys if the attribute is not mentioned.

User query: "{query}"
Example output: {"gender": "female", "color": "brown", "category": "scarf"}
"""

# 提取过滤条件
response = generate(model="llama3", prompt=prompt.format(query="Looking for a brown scarf for women"))
# 解析response得到filter_dict

4. 检索后二次校验

为确保结果完全匹配,可让Ollama对检索到的商品做二次校验:

# 假设已得到retrieved_docs
check_prompt = """
Check if the product matches the user query. Reply ONLY with "YES" or "NO".

User query: "{query}"
Product description: {product_desc}
"""

# 过滤不符合的商品
valid_docs = []
for doc in retrieved_docs:
    check_result = generate(model="llama3", prompt=check_prompt.format(query=query, product_desc=doc.page_content))
    if check_result.strip() == "YES":
        valid_docs.append(doc)

if not valid_docs:
    print("No matching products found.")

内容的提问来源于stack exchange,提问作者Advait Shendage

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 06:51:17