You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure Cognitive Search过滤失效:RAG场景下关键词过滤未生效

RAG系统关键词过滤失效问题及解决方案

问题背景

我搭建了RAG系统,尝试实现关键词/短语过滤,但添加Include过滤后效果不佳:

  • 文档是500-700token长度的片段,构成完整文件的11个片段(测试用例为简历)
  • 向量搜索正常,会返回5个相关片段,但设置过滤词为「Outlook」以仅获取微软Outlook相关项目时,返回结果仍包含其他项目内容
  • 这些片段会传入OpenAI Completion API,导致最终输出包含非目标内容

当前实现过滤的核心代码如下:

public SearchOptions? CreateSearchOptions(int searchTypeInt, 
                                        int k, 
                                        ReadOnlyMemory<float> embeddings, 
                                        ReadOnlyMemory<float> namedEntitiesEmbeddings, 
                                        string filter, 
                                        FilterAction filterAction)
{
    _logger.LogInformation("CreateSearchOptions entered");

    SearchOptions? searchOptions = null;
    try
    {
        SearchType searchType = (SearchType)searchTypeInt;

        System.FormattableString formattableStr = $"SegmentText ct '{filter}'";
        if (!String.IsNullOrWhiteSpace(filter))
        {
            if (filterAction == FilterAction.Include)
            {
                formattableStr = $"search.ismatch({filter}, 'SegmentText')";
            }
            else if (filterAction == FilterAction.Exclude)
            {
                formattableStr = $"NOT(search.ismatch({filter}, 'SegmentText'))";
            }
        }

        searchOptions = new SearchOptions
        {
            Size = k,
            Select = { "SegmentText", "NamedEntities", "docId", "segmentId", "Source", "TimeSrcModified", "TimeSrcCreated", "TimeIngested" },
            Filter = SearchFilter.Create(formattableStr) 
        };

        if ((searchType & SearchType.Vector) == SearchType.Vector)
        {
            searchOptions.VectorSearch = new VectorSearchOptions();
            VectorizedQuery vq = new VectorizedQuery(embeddings) { KNearestNeighborsCount = k, Fields = { "SegmentTextVector" } };
            searchOptions.VectorSearch.Queries.Add(vq);
            if (namedEntitiesEmbeddings.Length > 0)
            {
                vq = new VectorizedQuery(namedEntitiesEmbeddings) { KNearestNeighborsCount = k, Fields = { "SegmentNamedEntitiesVector" } };
                searchOptions.VectorSearch.Queries.Add(vq);
            }
        }
    }
    catch (Exception ex)
    {
        _logger.LogError(ex, ex.Message);
        return null;
    }

    return searchOptions;
}

解决方案(除prompt提示外)

1. 优化过滤语法,提升匹配精准度

当前使用的search.ismatch({filter}, 'SegmentText')可能存在匹配范围过宽的问题,可调整为更精准的匹配逻辑:

  • 精确短语匹配:如果需要严格匹配「Outlook」这个短语,修改过滤语句为:
    formattableStr = $"search.ismatch('\"{filter}\"', 'SegmentText', 'full', 'any')";
    
    这里的"\"{filter}\""会将过滤词作为精确短语匹配,'full'指定全词匹配模式,避免部分字符匹配的干扰。
  • 直接包含检查:如果只需要片段包含过滤词即可,改用contains函数更直接可靠:
    formattableStr = $"contains(SegmentText, '{filter}')";
    

2. 内存二次过滤,确保结果合规

在向量搜索返回结果后,不要直接传入LLM,而是在内存中对结果做二次校验:

  • 遍历每个返回的片段,检查SegmentText字段是否包含目标过滤词
  • 只保留符合条件的片段,再将这些片段传入OpenAI Completion API
  • 示例逻辑(伪代码):
    var searchResults = await _searchClient.SearchAsync<DocumentSegment>(searchOptions);
    var filteredResults = searchResults.Value.Where(r => r.Document.SegmentText.Contains(filter, StringComparison.OrdinalIgnoreCase)).ToList();
    // 将filteredResults传入LLM
    

3. 调整文档拆分策略,降低片段主题混杂

当前500-700token的片段可能包含多个主题(比如同时有Outlook和其他项目内容),导致过滤后仍混有非目标内容:

  • 改为按语义单元拆分文档,比如按项目、经历段落拆分,每个片段只对应单一主题
  • 拆分时可借助NLP工具识别段落边界或主题切换点,确保片段主题聚焦
  • 这样过滤后返回的片段只会包含Outlook相关内容,不会混杂其他项目

4. 利用命名实体字段做精准过滤

代码中已包含NamedEntities字段,若「Outlook」已被预提取为命名实体,可直接针对该字段过滤:

  • 修改过滤条件为检查NamedEntities是否包含目标词,比全文匹配更精准:
    formattableStr = $"contains(NamedEntities, '{filter}')";
    
  • 若命名实体是数组格式,可调整为对应的过滤语法(比如Azure Cognitive Search中的any操作符):
    formattableStr = $"NamedEntities/any(ne: eq(ne, '{filter}'))";
    

内容的提问来源于stack exchange,提问作者Leon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 05:40:28