OpenSearch:如何按terms聚合桶统计字段的唯一值数量?
在Elasticsearch中按分组统计字段唯一值数量
完全可以实现,对应你在PostgreSQL里的COUNT(DISTINCT)分组查询需求,Elasticsearch需要通过嵌套聚合来完成:外层用terms聚合按年份分组,内层用cardinality聚合统计每组内颜色的唯一值数量。
具体查询DSL
{ "size": 0, "aggs": { "group_by_year": { "terms": { "field": "year" }, "aggs": { "distinct_color_count": { "cardinality": { "field": "color" } } } } } }
代码解释
size: 0:不需要返回具体文档,只关注聚合结果group_by_year:外层terms聚合,按year字段的值分组,生成每个年份对应的桶distinct_color_count:内层cardinality聚合,在每个年份桶内计算color字段的唯一值数量,效果等价于SQL里的COUNT(DISTINCT color)
补充说明
cardinality聚合是近似统计(基于HyperLogLog算法),默认精度下误差很小,适合大数据量场景。如果需要精确统计(数据量较小时适用),可以把内层换成terms聚合统计颜色的桶数,但性能会随颜色基数增长而下降:
{ "size": 0, "aggs": { "group_by_year": { "terms": { "field": "year" }, "aggs": { "color_terms": { "terms": { "field": "color", "size": 10000 // 设足够大的值覆盖所有颜色 } }, "distinct_color_count": { "bucket_script": { "buckets_path": { "colorCount": "color_terms._bucket_count" }, "script": "params.colorCount" } } } } } }
内容的提问来源于stack exchange,提问作者kingkupps
相关产品推荐
相关产品推荐

