You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Elasticsearch 6中按userName字段查找非分词重复文档的DSL写法

解决Elasticsearch 6中按完整字段值聚合查找重复文档的问题

问题根源

你遇到的中文分词拆分问题,本质是userName字段默认作为text类型会被中文分词器拆分,直接用该字段聚合时,会按分词后的单字/词组分组,而非完整字段值。要实现类似SQL的group by效果,必须用不分词的字段做聚合。

解决方案

1. 优先使用字段的keyword子字段(推荐)

如果userName是默认动态映射生成的,Elasticsearch 6会自动为text类型字段生成字段名.keyword格式的子字段,该子字段不会被分词,直接存储完整字符串。对应的DSL如下:

GET /es_index/_search
{
  "size": 0, // 无需返回具体文档,仅获取聚合结果
  "aggs": {
    "duplicate_user_names": {
      "terms": {
        "field": "userName.keyword", // 使用不分词的keyword子字段
        "min_doc_count": 2 // 对应SQL的having count(*) > 1
      }
    }
  }
}

执行后会得到符合预期的结果:

{
  "aggregations": {
    "duplicate_user_names": {
      "buckets": [
        {
          "key": "上海某公司",
          "doc_count": 2
        }
      ]
    }
  }
}

2. 无keyword子字段时,用脚本聚合(临时方案)

如果字段未配置keyword子字段且暂时无法修改映射,可以通过脚本直接读取字段原始值聚合(性能弱于keyword字段,数据量大时不推荐):

GET /es_index/_search
{
  "size": 0,
  "aggs": {
    "duplicate_user_names": {
      "terms": {
        "script": {
          "source": "doc['userName'].value" // 读取字段原始值
        },
        "min_doc_count": 2
      }
    }
  }
}

3. 永久解决:修改字段映射添加keyword子字段

若需长期使用该聚合场景,建议修改字段映射并重新索引数据:

  1. 关闭索引:
POST /es_index/_close
  1. 更新映射:
PUT /es_index/_mapping
{
  "properties": {
    "userName": {
      "type": "text",
      "fields": {
        "keyword": {
          "type": "keyword",
          "ignore_above": 256 // 限制字段长度,超出部分会被忽略
        }
      }
    }
  }
}
  1. 重新打开索引:
POST /es_index/_open
  1. 重新索引已有数据(使新映射生效):
POST _reindex
{
  "source": {
    "index": "es_index"
  },
  "dest": {
    "index": "es_index_new"
  }
}

(后续可删除原索引,将新索引别名改为es_index,或直接覆盖原索引,根据实际情况操作)

内容的提问来源于stack exchange,提问作者Joshua

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 14:12:39