You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在TypeScript中获取字符串中出现频率最高的N个单词?

从带标点文本中提取频率最高的N个单词(JS/TS实现)

步骤拆解

1. 文本预处理

  • 把所有字母转成小写,统一大小写规则
  • 过滤掉除了字母和单词内部撇号之外的所有标点符号(比如句号、逗号、感叹号等),保留it's、mom's这类带撇号的完整单词

2. 统计单词频率

  • 将处理后的文本分割成单词数组
  • 用Map统计每个单词的出现次数

3. 排序并取前N个

  • 把统计结果转换成数组,按频率从高到低排序
  • 截取前N个结果

完整代码示例

// 示例文本
const sampleText = "hello world this is taco here is some foo bar text to say hello to my world of tacos in the world of text and it is very cool thanks stackoverflow for it's my birthday. This text also contains punctuation and my mom's car and periods and such. I like apples, pie, and apple pie. Case should be ignored so case and Case are the same. It's and its are two different words!";

function getTopNWords(text: string, n: number): { word: string; count: number }[] {
  // 1. 预处理:转小写,匹配合法单词(字母+内部撇号)
  const lowerText = text.toLowerCase();
  // 正则解释:匹配以字母开头,可包含字母和内部撇号,结尾为字母的单词
  const words = lowerText.match(/\b[a-z]+(?:['’][a-z]+)?\b/g) || [];

  // 2. 统计频率
  const frequencyMap = new Map<string, number>();
  for (const word of words) {
    frequencyMap.set(word, (frequencyMap.get(word) || 0) + 1);
  }

  // 3. 排序并取前N个
  const sortedWords = Array.from(frequencyMap.entries())
    .sort((a, b) => b[1] - a[1])
    .slice(0, n)
    .map(([word, count]) => ({ word, count }));

  return sortedWords;
}

// 测试:取频率最高的5个单词
const top5Words = getTopNWords(sampleText, 5);
console.log(top5Words);

代码说明

  • 正则匹配:/\b[a-z]+(?:['’][a-z]+)?\b/g 确保只提取合法单词,it's会被保留,单独的标点或无效字符会被过滤
  • 大小写处理:转小写后统计,保证Case和case被视为同一个单词
  • 频率统计:用Map比普通对象更安全,避免原型链冲突
  • 排序逻辑:按频率降序排列,频率相同的单词会保持原出现顺序

运行结果

对于示例文本,运行后输出的前5个单词是:

[
  { word: 'and', count: 5 },
  { word: 'is', count: 3 },
  { word: 'world', count: 3 },
  { word: 'text', count: 3 },
  { word: 'my', count: 2 }
]

内容的提问来源于stack exchange,提问作者xXx_emo_girl_xXx

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 20:55:23