You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

按相似度对Significant Terms聚合返回的词项进行分组

对Elasticsearch Significant Terms聚合结果进行同义词分组处理

需求说明

需要将Significant Terms聚合返回的同义词或相似词项合并分组,例如把"ok"和"okay"合并为一组,汇总doc_count和bg_count,保留组内最高的score。

原始聚合结果:

[
  {
    "key" : "ok",
    "doc_count" : 200,
    "score" : 8.583258052060206E-4,
    "bg_count" : 213
  },
  {
    "key" : "okay",
    "doc_count" : 117,
    "score" : 4.814546694690713E-4,
    "bg_count" : 126
  },
  {
    "key" : "something else",
    "doc_count" : 100,
    "score" : 2.3240213379936128E-4,
    "bg_count" : 78
  }
]

目标分组结果:

[
  {
    "grouped_keys" : ["ok","okay"],
    "doc_count" : 317,
    "score" : 8.583258052060206E-4,
    "bg_count" : 339
  },
  {
    "grouped_keys" : ["something else"],
    "doc_count" : 100,
    "score" : 2.3240213379936128E-4,
    "bg_count" : 78
  }
]

解决方案

方案1:索引阶段配置同义词过滤器(推荐)

在创建索引时,通过同义词分析器让Elasticsearch在分词阶段就合并同义词,这样Significant Terms聚合会直接返回分组后的结果。

  1. 创建带同义词分析器的索引:
PUT /my_index
{
  "settings": {
    "analysis": {
      "filter": {
        "synonym_filter": {
          "type": "synonym",
          "synonyms": ["ok, okay"]
        }
      },
      "analyzer": {
        "synonym_analyzer": {
          "tokenizer": "standard",
          "filter": ["lowercase", "synonym_filter"]
        }
      }
    }
  },
  "mappings": {
    "properties": {
      "content": {
        "type": "text",
        "analyzer": "synonym_analyzer"
      }
    }
  }
}
  1. 执行Significant Terms聚合时,使用配置了同义词分析的字段:
POST /my_index/_search
{
  "size": 0,
  "aggs": {
    "significant_terms": {
      "significant_terms": {
        "field": "content"
      }
    }
  }
}

这样聚合结果会直接将"ok"和"okay"合并为一个词项,无需后续处理。

方案2:聚合结果后处理(无索引修改权限时使用)

如果无法修改现有索引配置,可以在获取聚合结果后,通过代码手动分组合并。以下是Python示例:

# 原始Significant Terms聚合结果
original_results = [
  {"key": "ok", "doc_count": 200, "score": 8.583258052060206E-4, "bg_count": 213},
  {"key": "okay", "doc_count": 117, "score": 4.814546694690713E-4, "bg_count": 126},
  {"key": "something else", "doc_count": 100, "score": 2.3240213379936128E-4, "bg_count": 78}
]

# 定义同义词分组规则,可根据需求扩展
synonym_groups = [["ok", "okay"], ["something else"]]

grouped_results = []
for group_keys in synonym_groups:
    # 筛选当前组的所有结果项
    group_items = [item for item in original_results if item["key"] in group_keys]
    if not group_items:
        continue
    # 计算合并后的字段值
    grouped_results.append({
        "grouped_keys": group_keys,
        "doc_count": sum(item["doc_count"] for item in group_items),
        "score": max(item["score"] for item in group_items),
        "bg_count": sum(item["bg_count"] for item in group_items)
    })

print(grouped_results)

方案3:动态相似词分组(复杂场景)

如果需要动态识别语义相似的词项(而非固定同义词),可以使用Elasticsearch的ELSER模型生成词项的语义向量,然后通过向量相似度聚类(如K-means)分组,最后再聚合统计。这种方式需要额外配置ELSER模型,适合需要自动识别相似词的场景。


内容的提问来源于stack exchange,提问作者Christian Dove

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 20:43:26