如何在Elasticsearch的query_string搜索中聚合匹配的通配符词条?
问题需求
我需要在嵌套字典列表中搜索带通配符的词条,然后按匹配的通配符分组,获取对应的词条列表及uuid信息。
索引结构
我的索引settings和mapping如下:
{ "settings": { "analysis": { "char_filter": { "my_filter": { "type": "mapping", "mappings": [ "- => _", ] }, }, "analyzer": { "my_analyzer": { "tokenizer": "standard", "char_filter": [ "my_filter" ], "filter": [ "lowercase", ] } } } }, "mappings": { "properties": { "uuid": { "type": "keyword" }, "urls": { "type": "nested", "properties": { "url": { "type": "keyword" }, "is_visited": { "type": "boolean" } } } } } }
数据示例
存储的数据样例:
{ "uuid":"afa9ac03-0723-4d66-ae18-08a51e2973bd", "urls": [ { "is_visited": true, "url": "https://www.google.com" }, { "is_visited": false, "url": "https://www.facebook.com" }, { "is_visited": true, "url": "https://www.twitter.com" } ] }, { "uuid":"4a1c695d-756b-4d9d-b3a0-cf524d955884", "urls": [ { "is_visited": true, "url": "https://www.stackoverflow.com" }, { "is_visited": false, "url": "https://www.facebook.com" }, { "is_visited": false, "url": "https://drive.google.com" }, { "is_visited": false, "url": "https://maps.google.com" } ] }
期望结果
通过通配符查询"*google.com OR *twitter.com",得到按通配符分组的聚合结果:
"hits": { "*google.com": [ { "uuid": "4a1c695d-756b-4d9d-b3a0-cf524d955884", "_source": { "is_visited": false, "url": "https://drive.google.com" } }, { "uuid": "4a1c695d-756b-4d9d-b3a0-cf524d955884", "_source": { "is_visited": false, "url": "https://maps.google.com" } }, { "uuid":"afa9ac03-0723-4d66-ae18-08a51e2973bd", "_source": { "is_visited": true, "url": "https://www.google.com" } } ], "*twitter.com": [ { "uuid":"afa9ac03-0723-4d66-ae18-08a51e2973bd", "_source": { "is_visited": true, "url": "https://www.twitter.com" } } ] }
当前问题
我用以下Python查询,返回的是每个匹配词条的单独命中,不是按通配符分组的格式:
body = { #"_source": False, "size": 100, "query": { "nested": { "path": "urls", "query":{ "query_string":{ "query": f"urls.url:{urlToSearch}", } } ,"inner_hits": { "size":100 # returns top 100 results } } } }
解决方案
Elasticsearch原生无法直接返回你期望的分组格式,需要通过嵌套聚合+客户端结果处理来实现。核心思路是:针对每个通配符条件做嵌套聚合,再把聚合结果整理成目标格式。
步骤1:构造带聚合的查询
针对每个通配符规则,创建单独的嵌套聚合,同时保留匹配的url和uuid信息:
url_patterns = ["*google.com", "*twitter.com"] body = { "size": 0, # 不需要返回顶层文档,只看聚合结果 "aggs": { "group_by_pattern": { "filters": { "filters": { pattern: { "nested": { "path": "urls", "query": { "wildcard": { "urls.url": pattern } } } } for pattern in url_patterns } }, "aggs": { "nested_urls": { "nested": { "path": "urls" }, "aggs": { "match_urls": { "filter": { "bool": { "should": [ {"wildcard": {"urls.url": pattern}} for pattern in url_patterns ] } }, "aggs": { "top_url_hits": { "top_hits": { "size": 100, "_source": ["urls.url", "urls.is_visited"] } }, # 关联顶层文档的uuid "parent_uuid": { "reverse_nested": {}, "aggs": { "uuid": { "terms": { "field": "uuid", "size": 100 } } } } } } } } } } } }
步骤2:处理聚合结果
执行查询后,需要把ES返回的聚合数据整理成你期望的格式:
from elasticsearch import Elasticsearch es = Elasticsearch(["your_es_host:port"]) response = es.search(index="your_index_name", body=body) result = {} # 遍历每个通配符分组 for pattern, bucket in response["aggregations"]["group_by_pattern"]["buckets"].items(): result[pattern] = [] # 遍历每个匹配的嵌套url for hit in bucket["nested_urls"]["match_urls"]["top_url_hits"]["hits"]["hits"]: # 获取对应的uuid uuid = bucket["nested_urls"]["match_urls"]["parent_uuid"]["uuid"]["buckets"][0]["key"] result[pattern].append({ "uuid": uuid, "_source": hit["_source"] }) # 输出最终结果 print(result)
关键说明
- filters聚合:用来按通配符规则分组,每个规则对应一个桶。
- 嵌套+反向嵌套聚合:既要匹配嵌套的url字段,又要关联顶层文档的uuid。
- top_hits聚合:保留匹配的url具体信息(is_visited和url值)。
- 客户端处理:ES返回的聚合结构比较层级化,需要手动提取并组装成目标格式。
如果需要处理大量数据,可以调整size参数,或者使用滚动查询避免内存溢出。
内容的提问来源于stack exchange,提问作者Sandy
相关产品推荐
相关产品推荐

