Azure Cognitive Search过滤失效:RAG场景下关键词过滤未生效
RAG系统关键词过滤失效问题及解决方案
问题背景
我搭建了RAG系统,尝试实现关键词/短语过滤,但添加Include过滤后效果不佳:
- 文档是500-700token长度的片段,构成完整文件的11个片段(测试用例为简历)
- 向量搜索正常,会返回5个相关片段,但设置过滤词为「Outlook」以仅获取微软Outlook相关项目时,返回结果仍包含其他项目内容
- 这些片段会传入OpenAI Completion API,导致最终输出包含非目标内容
当前实现过滤的核心代码如下:
public SearchOptions? CreateSearchOptions(int searchTypeInt, int k, ReadOnlyMemory<float> embeddings, ReadOnlyMemory<float> namedEntitiesEmbeddings, string filter, FilterAction filterAction) { _logger.LogInformation("CreateSearchOptions entered"); SearchOptions? searchOptions = null; try { SearchType searchType = (SearchType)searchTypeInt; System.FormattableString formattableStr = $"SegmentText ct '{filter}'"; if (!String.IsNullOrWhiteSpace(filter)) { if (filterAction == FilterAction.Include) { formattableStr = $"search.ismatch({filter}, 'SegmentText')"; } else if (filterAction == FilterAction.Exclude) { formattableStr = $"NOT(search.ismatch({filter}, 'SegmentText'))"; } } searchOptions = new SearchOptions { Size = k, Select = { "SegmentText", "NamedEntities", "docId", "segmentId", "Source", "TimeSrcModified", "TimeSrcCreated", "TimeIngested" }, Filter = SearchFilter.Create(formattableStr) }; if ((searchType & SearchType.Vector) == SearchType.Vector) { searchOptions.VectorSearch = new VectorSearchOptions(); VectorizedQuery vq = new VectorizedQuery(embeddings) { KNearestNeighborsCount = k, Fields = { "SegmentTextVector" } }; searchOptions.VectorSearch.Queries.Add(vq); if (namedEntitiesEmbeddings.Length > 0) { vq = new VectorizedQuery(namedEntitiesEmbeddings) { KNearestNeighborsCount = k, Fields = { "SegmentNamedEntitiesVector" } }; searchOptions.VectorSearch.Queries.Add(vq); } } } catch (Exception ex) { _logger.LogError(ex, ex.Message); return null; } return searchOptions; }
解决方案(除prompt提示外)
1. 优化过滤语法,提升匹配精准度
当前使用的search.ismatch({filter}, 'SegmentText')可能存在匹配范围过宽的问题,可调整为更精准的匹配逻辑:
- 精确短语匹配:如果需要严格匹配「Outlook」这个短语,修改过滤语句为:
这里的formattableStr = $"search.ismatch('\"{filter}\"', 'SegmentText', 'full', 'any')";"\"{filter}\""会将过滤词作为精确短语匹配,'full'指定全词匹配模式,避免部分字符匹配的干扰。 - 直接包含检查:如果只需要片段包含过滤词即可,改用
contains函数更直接可靠:formattableStr = $"contains(SegmentText, '{filter}')";
2. 内存二次过滤,确保结果合规
在向量搜索返回结果后,不要直接传入LLM,而是在内存中对结果做二次校验:
- 遍历每个返回的片段,检查
SegmentText字段是否包含目标过滤词 - 只保留符合条件的片段,再将这些片段传入OpenAI Completion API
- 示例逻辑(伪代码):
var searchResults = await _searchClient.SearchAsync<DocumentSegment>(searchOptions); var filteredResults = searchResults.Value.Where(r => r.Document.SegmentText.Contains(filter, StringComparison.OrdinalIgnoreCase)).ToList(); // 将filteredResults传入LLM
3. 调整文档拆分策略,降低片段主题混杂
当前500-700token的片段可能包含多个主题(比如同时有Outlook和其他项目内容),导致过滤后仍混有非目标内容:
- 改为按语义单元拆分文档,比如按项目、经历段落拆分,每个片段只对应单一主题
- 拆分时可借助NLP工具识别段落边界或主题切换点,确保片段主题聚焦
- 这样过滤后返回的片段只会包含Outlook相关内容,不会混杂其他项目
4. 利用命名实体字段做精准过滤
代码中已包含NamedEntities字段,若「Outlook」已被预提取为命名实体,可直接针对该字段过滤:
- 修改过滤条件为检查
NamedEntities是否包含目标词,比全文匹配更精准:formattableStr = $"contains(NamedEntities, '{filter}')"; - 若命名实体是数组格式,可调整为对应的过滤语法(比如Azure Cognitive Search中的
any操作符):formattableStr = $"NamedEntities/any(ne: eq(ne, '{filter}'))";
内容的提问来源于stack exchange,提问作者Leon
相关产品推荐
相关产品推荐

