如何在Elasticsearch中统计逗号分隔interests字段各值的文档出现次数
Elasticsearch 逗号分隔字段值统计实现步骤
我们提供两种方案适配不同使用场景:
方案1:临时查询(无需修改索引结构)
适用于偶发查询场景,不需要调整现有索引配置,用运行时字段实时拆分统计:
- 查询时先定义运行时字段,用脚本将
interests字段按逗号拆分为数组,再通过terms聚合统计每个值的文档数 - 执行查询DSL如下:
GET 替换为你的索引名/_search { "size": 0, "runtime_mappings": { "interests_split": { "type": "keyword", "script": { "source": """ if(doc['interests'].size() == 0) return; String[] arr = doc['interests'].value.splitOnToken(','); for(String s : arr) { emit(s); } """ } } }, "aggs": { "interest_count": { "terms": { "field": "interests_split", "size": 100 // 可根据实际分类数量调整上限 } } } }
- 返回结果中
aggregations.interest_count.buckets即为你需要的统计结果,结构和你预期的完全一致。
如果你的Elasticsearch版本低于7.11不支持运行时字段,可以直接在terms聚合中内嵌脚本实现:
GET 替换为你的索引名/_search { "size": 0, "aggs": { "interest_count": { "terms": { "script": { "source": "doc['interests'].value.splitOnToken(',')" }, "size": 100 } } } }
方案2:高频查询场景(修改索引结构优化性能)
如果该统计是高频需求,运行时字段每次实时计算性能较差,可通过调整索引配置预拆分字段:
- 给索引新增自定义逗号分词器,索引数据时自动拆分
interests字段 - 修改
interests字段mapping,新增拆分后的子字段 - 对历史数据执行更新让配置生效
- 直接对子字段做terms聚合即可得到统计结果
完整操作命令:
// 1. 关闭索引(仅已有索引需要,新建索引跳过) POST 替换为你的索引名/_close // 2. 新增逗号分隔分词器(新建索引写到settings配置中即可) PUT 替换为你的索引名/_settings { "analysis": { "analyzer": { "comma_split": { "type": "pattern", "pattern": ",", "lowercase": false // 不需要统一大小写可保持关闭 } } } } // 3. 重新开启索引(仅已有索引需要,新建索引跳过) POST 替换为你的索引名/_open // 4. 修改字段mapping PUT 替换为你的索引名/_mapping { "properties": { "interests": { "type": "keyword", "fields": { "split": { "type": "text", "analyzer": "comma_split", "fielddata": true } } } } } // 5. 历史数据更新生效(仅已有索引需要) POST 替换为你的索引名/_update_by_query?conflicts=proceed
后续查询DSL:
GET 替换为你的索引名/_search { "size": 0, "aggs": { "interest_count": { "terms": { "field": "interests.split", "size": 100 } } } }
内容的提问来源于stack exchange,提问作者fajri
相关产品推荐
相关产品推荐

