Elasticsearch如何排除跨句子的搜索结果?R语言elastic库实操问题
解决Elasticsearch特定间隔查询并排除含句号文档的问题
问题背景
使用R语言elastic库向Elasticsearch的test索引导入了12个文档,需求是实现:
dog MUST PRECEDENT that, WITH A MAX_GAP OF 3, BUT MUST NOT INCLUDE .
即仅返回最后一个无句号的文档:"is THIS a brown dog apple that likes to be."
现有尝试的问题
- 最初使用默认分析器时,Elasticsearch默认不索引句号,导致
must_not匹配.的规则无效,返回全部12个文档。 - 改用
whitespace分析器后,返回了两个文档:目标文档和"THIS a brown dog is. that likes to be."。原因是该文档的句号包含在is.中,不在dog与that的间隔范围内,must_not规则无法过滤掉它。
解决方案
需要从索引映射和查询逻辑两方面调整:
1. 自定义分析器,精准拆分标点与单词
使用pattern分析器自定义分词规则,将标点作为独立token拆分,确保句号能被正确识别:
index_create(connection, indexName) mapping <- ' { "properties": { "full_text": { "type": "text", "analyzer": "custom_punct_analyzer" } }, "settings": { "analysis": { "analyzer": { "custom_punct_analyzer": { "tokenizer": "custom_punct_tokenizer", "filter": ["lowercase"] } }, "tokenizer": { "custom_punct_tokenizer": { "type": "pattern", "pattern": "(\\s+|\\p{Punct})", "group": -1 } } } } } ' mapping_create(connection, indexName, body = mapping)
该分析器会将空格和标点作为分隔符,把is.拆分为is和.两个独立token,同时支持小写转换(不需要可移除lowercase过滤器)。
2. 优化查询逻辑,双重限制间隔与全局排除
调整查询规则,既确保dog与that的间隔内无句号,又全局排除所有含句号的文档:
query <- '{ "from": 0, "size": 10000, "query": { "bool": { "must": [ { "intervals": { "full_text": { "all_of": { "ordered": true, "intervals": [ { "match": { "query": "dog", "max_gaps": 0 } }, { "match": { "query": "that", "max_gaps": 0 } } ], "max_gaps": 3, "filter": { "not": { "match": { "query": ".", "max_gaps": 0 } } } } } } } ], "must_not": [ { "match": { "full_text": "." } } ] } } }' Search(connection, indexName, body = query)
关键调整:
- 给
all_of添加filter规则,确保dog与that的间隔区间内不出现句号 - 保留全局
must_not规则排除所有含句号的文档,实现双重校验 - 简化原查询中冗余的
any_of逻辑,直接匹配that更贴合需求
通过以上调整,即可精准返回唯一符合条件的文档。
内容的提问来源于stack exchange,提问作者Pxu80
相关产品推荐
相关产品推荐

