You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在MongoDB中实现文本字段词频统计,获取TOP100高频词

MongoDB Shell 词频统计实现方案

以下聚合查询和你给出的Python代码功能完全等效,假设你的数据集存储在名为elonmusk的集合中,待统计的文本字段为tweet,如果你的集合名、待统计字段名和示例不同,替换代码中对应的名称即可:

db.elonmusk.aggregate([
  // 拆分tweet文本为单词数组
  {
    $project: {
      words: { $split: ["$tweet", " "] }
    }
  },
  // 展开单词数组为单单词文档
  { $unwind: "$words" },
  // 按单词分组统计出现次数
  {
    $group: {
      _id: "$words",
      count: { $sum: 1 }
    }
  },
  // 按出现次数倒序排序
  { $sort: { count: -1 } },
  // 取前100个高频词
  { $limit: 100 }
])

返回结果的_id字段为对应单词,count字段为该单词的出现总次数。

优化版本(解决大小写、标点干扰问题)

如果需要统一单词大小写、移除附着在单词上的标点,避免同类单词被拆分统计,可以用以下版本:

db.elonmusk.aggregate([
  {
    $project: {
      // 先转全小写,再移除常见标点,最后按空格拆分单词
      words: {
        $split: [
          {
            $replaceAll: {
              input: { $toLower: "$tweet" },
              find: /[.,!?;:'"]/g,
              replacement: ""
            }
          },
          " "
        ]
      }
    }
  },
  // 展开数组并过滤空字符串
  { $unwind: "$words" },
  { $match: { words: { $ne: "" } } },
  // 分组统计+排序+取前100
  {
    $group: {
      _id: "$words",
      count: { $sum: 1 }
    }
  },
  { $sort: { count: -1 } },
  { $limit: 100 }
])

内容的提问来源于stack exchange,提问作者Christopher

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 06:57:02