You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

ElasticSearch 7.15 如何统计单文档单字段的高频词及出现次数

问题根因
  • 你当前使用的pdf_content字段为keyword类型,ES不会对该类型字段的内容做分词,整段字符串会被识别为1个独立的term,无法拆分出单个词汇统计频次
  • 你编写的terms聚合逻辑是统计全索引中不同term对应的关联文档数量,并非统计单个文档内部的词汇出现次数,且未添加过滤条件定位到目标文档,因此无法得到预期结果
实现方案

根据你是否可以调整索引mapping,可选择以下两种方案:

方案1:无需修改索引(临时查询适用)

直接通过runtime字段在查询时动态拆分内容做聚合,查询语句如下,将<目标文档ID>替换为你要统计的文档的_id即可:

GET /pdf/_search
{
  "size": 0,
  "query": {
    "ids": {
      "values": ["<目标文档ID>"]
    }
  },
  "runtime_mappings": {
    "split_pdf_word": {
      "type": "keyword",
      "script": "for (String word : doc['pdf_content'].value.split(' ')) { emit(word); }"
    }
  },
  "aggs": {
    "word_frequency": {
      "terms": {
        "field": "split_pdf_word",
        "order": {
          "_count": "desc"
        }
      }
    }
  }
}

返回结果的aggregations.word_frequency.buckets中,key对应词汇,doc_count对应出现频次,和你预期的结构完全匹配。如果你的内容不是纯空格分隔,可将split参数替换为对应的正则表达式,比如split(/\W+/)可按任意非单词字符拆分。

方案2:调整索引mapping(高频查询适用)

如果需要经常做这类统计,建议调整索引结构新增text类型子字段,性能更高:

1. 更新mapping

PUT /pdf/_mapping
{
  "properties": {
    "pdf_content": {
      "type": "keyword",
      "fields": {
        "text": {
          "type": "text",
          "analyzer": "standard"
        }
      }
    }
  }
}

2. 刷新存量数据让mapping生效

POST /pdf/_update_by_query?conflicts=proceed

3. 用词向量接口查询单个文档词频

GET /pdf/_termvectors/<目标文档ID>
{
  "fields": ["pdf_content.text"],
  "positions": false,
  "offsets": false,
  "field_statistics": false
}

返回结果的terms字段中每个词汇的term_freq即为出现频次。

内容的提问来源于stack exchange,提问作者propre_poli

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 22:06:03