Elasticsearch过滤条件未按预期生效,must_not未排除指定文档
Elasticsearch must_not过滤条件未生效问题解决
问题背景
文档结构
{ "_index": "terms", "_type": "_doc", "_id": "30028861", "_version": 36, "_seq_no": 338988, "_primary_term": 6, "found": true, "_source": { "title": { "value": "Credit Card Number", "is_changed": false }, "node_id": 30028861, "domain_id": 30028838, "node_name": "Credit Card Number", "parent_id": 30028839, "term_name": "Credit Card Number", "account_id": 30028807, "published_status": "Published", "published_status_id": "PUB", "node_sub_type": "TRM" } }
查询需求与问题
需要筛选满足以下条件的文档:
account_id为30028807domain_id在[30119066, 30123338]或parent_id为30028839published_status_id不为"PUB"且published_status不为"Published"published_status.raw不为"Disabled"
但编写的查询语句返回结果仍包含published_status_id: PUB和published_status: Published的文档,过滤未生效。
原查询语句
{ "from":0, "size":10, "track_total_hits":true, "query":{ "bool":{ "must":[ { "match":{ "account_id":30028807 } } ], "filter":[ { "bool":{ "should":[ { "terms":{ "domain_id":[ 30119066, 30123338 ] } }, { "terms":{ "parent_id":[ 30028839 ] } } ] } }, { "bool":{ "must_not":[ { "terms":{ "published_status_id":[ "PUB" ] } }, { "terms":{ "published_status":[ "Published" ] } } ] } } ], "must_not":{ "terms":{ "published_status.raw":[ "Disabled" ] } } } }, "_source":[ "node_type", "node_id", "parent_name", "table_name" ] }
问题原因
- 字段匹配不精确:若
published_status是text类型且使用默认分词器,存储时会被转为小写(如"published"),原查询中用terms精确匹配首字母大写的"Published",无法匹配到实际存储的值,导致must_not条件失效。 - 查询冗余:
published_status_id和published_status是一一对应关系(PUB对应Published),同时排除两个字段属于冗余操作,且可能因其中一个字段匹配失败导致整体过滤失效。
修正后的查询语句
{ "from": 0, "size": 10, "track_total_hits": true, "query": { "bool": { "must": [ { "match": { "account_id": 30028807 } } ], "filter": [ { "bool": { "should": [ { "terms": { "domain_id": [30119066, 30123338] } }, { "terms": { "parent_id": [30028839] } } ], "minimum_should_match": 1 } }, { "bool": { "must_not": [ { "term": { "published_status_id": "PUB" } } ] } } ], "must_not": { "term": { "published_status.raw": "Disabled" } } } }, "_source": ["node_type", "node_id", "parent_name", "table_name"] }
关键调整点
- 用
term替代terms处理单个值查询,提升效率 - 仅保留对
published_status_id: PUB的排除(因与published_status: Published一一对应),简化逻辑 - 若需严格同时排除两个字段,需使用
published_status的keyword子字段进行精确匹配,示例如下:
{ "from": 0, "size": 10, "track_total_hits": true, "query": { "bool": { "must": [ { "match": { "account_id": 30028807 } } ], "filter": [ { "bool": { "should": [ { "terms": { "domain_id": [30119066, 30123338] } }, { "terms": { "parent_id": [30028839] } } ] } }, { "bool": { "must_not": [ { "term": { "published_status_id": "PUB" } }, { "term": { "published_status.keyword": "Published" } } ] } } ], "must_not": { "term": { "published_status.raw": "Disabled" } } } }, "_source": ["node_type", "node_id", "parent_name", "table_name"] }
内容的提问来源于stack exchange,提问作者Ninja
相关产品推荐
相关产品推荐

