You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在OpenSearch索引中统计唯一词汇对应的文档数量

OpenSearch统计词汇对应文档数量的查询方案

前提准备

首先确保你的文本字段(假设字段名为content)支持聚合,因为OpenSearch的text类型字段默认禁用了fielddata,需要先更新索引映射开启:

PUT test_index/_mapping
{
  "properties": {
    "content": {
      "type": "text",
      "fielddata": true
    }
  }
}

查询语句

使用terms聚合来统计每个词汇的文档出现次数,同时设置size参数确保能返回所有唯一词汇(根据实际词汇量调整):

GET test_index/_search
{
  "size": 0, // 不返回原始文档,只返回聚合结果
  "aggs": {
    "word_document_counts": {
      "terms": {
        "field": "content",
        "size": 1000, // 调整为你需要的最大词汇数量
        "order": {
          "_count": "desc" // 按文档数从多到少排序
        }
      }
    }
  }
}

结果说明

返回的聚合结果中,aggregations.word_document_counts.buckets数组里的每个元素对应一个词汇:

  • key:唯一词汇(比如brown、fox)
  • doc_count:该词汇出现的文档数量(比如brown对应2,quick对应1)

优化建议

如果需要统一大小写(避免The和the被统计为两个不同词汇),可以给字段添加小写分词的子字段:

PUT test_index/_mapping
{
  "properties": {
    "content": {
      "type": "text",
      "fielddata": true,
      "fields": {
        "lowercase": {
          "type": "text",
          "fielddata": true,
          "analyzer": "lowercase"
        }
      }
    }
  }
}

查询时改用content.lowercase字段即可:

GET test_index/_search
{
  "size": 0,
  "aggs": {
    "word_document_counts": {
      "terms": {
        "field": "content.lowercase",
        "size": 1000,
        "order": {
          "_count": "desc"
        }
      }
    }
  }
}

内容的提问来源于stack exchange,提问作者Suvid Sahay

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 13:35:17