You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于FAISS与LangChain的食谱搜索系统如何实现排除式相似搜索?

问题

我基于FAISS向量存储、LangChain和Hugging Face的sentence-transformers/all-MiniLM-L6-v2嵌入模型构建食谱搜索系统,支持按食材、饮食限制、菜系、餐型等参数查找食谱。我用以下Prompt生成带饮食限制定义的自然语言查询和过滤术语:

query_prompt = ChatPromptTemplate.from_template(
    """Analyze the following search parameters for a recipe search query:
    Dietary Restrictions: {dietary_restrictions}    
    Ingredients: {ingredients}
    Dietary Restrictions: {dietary_restrictions}
    Cuisine: {cuisine}
    Meal Type: {meal_type}

    Transform these parameters into a natural language query that can be used to search for recipes. Include the definition of the dietary restrictions. 

    Provide your analysis in the following format:
    Query: [Transformed natural language query]
    Filter Terms: [List of key terms for filtering, including ingredients, dietary restrictions, cuisine, and meal type]
    """
)

生成的查询示例:"Dairy-free American snack recipes, excluding all products derived from milk such as cheese, yogurt, and butter."

随后执行相似搜索:

search_results = vector_store.similarity_search(query, k=1000)  # 先获取前1000条结果

但相似搜索无法正确排除指定术语,比如传入以下参数时:

find_recipes(
    ingredients=[],
    dietary_restrictions=["dairy free"],
    cuisine="American",
    meal_type="snacks"
)

返回结果仍包含含乳制品的食谱,比如“### 3. Milk Peda Recipe”“### 12. Yogurt Cheese (Labneh)”。

需要解决的核心问题:如何修改相似搜索以正确排除查询中‘excluding’部分的术语或概念?是否可实现负匹配?有哪些优化搜索以尊重排除项的思路?


解决方案

1. 先召回后硬过滤

利用Prompt生成的Filter Terms提取明确的排除术语,在相似搜索召回结果后做精准过滤,直接剔除包含禁用术语的食谱。

示例代码:

# 解析Prompt输出,提取排除术语(以 dairy-free 为例)
exclude_terms = ["cheese", "yogurt", "butter", "milk"]

# 执行初始相似搜索
search_results = vector_store.similarity_search(positive_query, k=1000)

# 硬过滤排除含禁用术语的文档
filtered_results = [
    doc for doc in search_results 
    if not any(term.lower() in doc.page_content.lower() for term in exclude_terms)
]

优点是实现简单、逻辑直观;缺点是若初始召回结果大多不符合要求,会浪费计算资源。

2. 负嵌入加权实现语义级负匹配

通过嵌入模型生成排除术语的向量,对查询嵌入做加权调整,从语义层面降低含排除概念食谱的相似度评分。

示例代码:

from langchain.embeddings import HuggingFaceEmbeddings
from sklearn.metrics.pairwise import cosine_similarity
import numpy as np

embeddings = HuggingFaceEmbeddings(model_name="sentence-transformers/all-MiniLM-L6-v2")

# 生成正查询嵌入(聚焦需要的食谱特征)
positive_query = "Dairy-free American snack recipes"
positive_embedding = embeddings.embed_query(positive_query)

# 生成排除术语的嵌入并取平均值
exclude_terms = ["cheese", "yogurt", "butter", "milk"]
exclude_embeddings = [embeddings.embed_query(term) for term in exclude_terms]
avg_exclude_embedding = np.mean(exclude_embeddings, axis=0)

# 调整查询嵌入:正嵌入减去加权后的负嵌入(权重可根据效果调整)
weight = 0.5
adjusted_query_embedding = np.array(positive_embedding) - weight * np.array(avg_exclude_embedding)

# 手动计算相似度并排序(适配FAISS自定义查询)
doc_embeddings = vector_store.index.reconstruct_n(0, vector_store.index.ntotal)
docs = list(vector_store.docstore._dict.values())

similarities = cosine_similarity([adjusted_query_embedding], doc_embeddings)[0]
sorted_indices = np.argsort(similarities)[::-1][:1000]
filtered_results = [docs[i] for i in sorted_indices]

这种方式能捕捉语义层面的排除关系,但需要反复调整权重参数,避免过度惩罚导致遗漏相关结果。

3. 优化Prompt明确正负匹配规则

修改Prompt,让输出清晰区分「需要匹配的正查询」和「需要排除的术语列表」,避免将排除项混在自然语言查询中,简化后续处理逻辑。

优化后的Prompt:

query_prompt = ChatPromptTemplate.from_template(
    """Analyze the following search parameters for a recipe search query:
    Dietary Restrictions: {dietary_restrictions}    
    Ingredients: {ingredients}
    Cuisine: {cuisine}
    Meal Type: {meal_type}

    1. Generate a natural language query focused on the recipes we WANT (include ingredients, dietary restrictions, cuisine, meal type).
    2. List all specific terms/ingredients we need to EXCLUDE based on dietary restrictions.

    Provide your analysis in the following format:
    Positive Query: [Natural language query for desired recipes]
    Exclude Terms: [List of exact terms/ingredients to exclude, e.g., cheese, yogurt, butter]
    """
)

生成的结果结构更清晰,可结合「正查询召回+排除术语硬过滤」或「负嵌入加权」的方式使用。

4. 结构化元数据过滤(若有条件)

如果食谱文档带有结构化元数据(如dietary_tags、ingredients_list字段),可直接利用FAISS的过滤功能,在搜索阶段就剔除不符合要求的结果。

示例代码:

# 假设文档元数据包含dietary_tags字段,标记了食谱的饮食属性
search_results = vector_store.similarity_search(
    query="Dairy-free American snack recipes",
    k=1000,
    filter={"dietary_tags": {"$in": ["dairy-free"]}}
)

这种方式效率最高、结果最准确,但依赖食谱数据的结构化程度。


内容的提问来源于stack exchange,提问作者Rasik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 21:20:08