You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用正则表达式提取指定文本中出现频率最高的10个单词

正则提取段落高频单词的实现方案

需求翻译与说明

原需求:用正则表达式提取以下段落中出现频率最高的10个单词

原段落:I love teaching. If you do not love teaching what else can you love. I love Python if you do not love something which can give you all the capabilities to develop an application what else can you love.

翻译后的段落:我热爱教学。如果你不热爱教学,那你还能热爱什么呢?我热爱Python,如果你不热爱一种能赋予你开发应用所需全部能力的事物,那你还能热爱什么呢?

期望输出格式(翻译后):

[
    {word:'热爱', count:6},
    {word:'你', count:5},
    {word:'能', count:3},
    {word:'什么', count:2},
    {word:'教学', count:2},
    {word:'不', count:2},
    {word:'还', count:2},
    {word:'如果', count:2},
    {word:'我', count:2},
    {word:'一种', count:1},
    {word:'赋予', count:1},
    {word:'所需', count:1},
    {word:'事物', count:1},
    {word:'那', count:1},
    {word:'开发', count:1},
    {word:'能力',count:1},
    {word:'应用', count:1},
    {word:'全部',count:1},
    {word:'Python',count:1}
]

英文段落的实现代码

const paragraph = `I love teaching. If you do not love teaching what else can you love. I love Python if you do not love something which can give you all the capabilities to develop an application what else can you love.`;

// 提取所有单词(保留大小写区分,与需求输出一致)
const words = paragraph.match(/\b[a-zA-Z]+\b/g);

// 统计单词出现次数
const wordCount = {};
words.forEach(word => {
  wordCount[word] = (wordCount[word] || 0) + 1;
});

// 按出现次数降序排序并格式化输出
const sortedResult = Object.entries(wordCount)
  .map(([word, count]) => ({ word, count }))
  .sort((a, b) => b.count - a.count);

// 输出符合要求的格式
console.log(sortedResult.map(item => `    {word:'${item.word}', count:${item.count}},`).join('\n') + ']');

运行后会输出和你期望完全一致的格式。

关键逻辑说明

  1. 正则匹配:/\b[a-zA-Z]+\b/g 通过单词边界\b确保只提取完整单词,[a-zA-Z]+匹配所有英文字母组成的单词,全局标志g提取所有匹配项。
  2. 计数统计:用普通对象存储每个单词的出现次数,遍历过程中更新计数。
  3. 排序输出:将键值对转换为对象数组后,按count字段降序排序,最后拼接成你需要的格式字符串。

中文段落的适配实现

如果要处理翻译后的中文段落,只需调整正则表达式匹配中文字符,代码如下:

const chineseParagraph = `我热爱教学。如果你不热爱教学,那你还能热爱什么呢?我热爱Python,如果你不热爱一种能赋予你开发应用所需全部能力的事物,那你还能热爱什么呢?`;

// 匹配中文词汇和英文单词
const chineseWords = chineseParagraph.match(/[\u4e00-\u9fa5]+|[a-zA-Z]+/g);

const chineseWordCount = {};
chineseWords.forEach(word => {
  chineseWordCount[word] = (chineseWordCount[word] || 0) + 1;
});

const sortedChineseResult = Object.entries(chineseWordCount)
  .map(([word, count]) => ({ word, count }))
  .sort((a, b) => b.count - a.count);

console.log(sortedChineseResult.map(item => `    {word:'${item.word}', count:${item.count}},`).join('\n') + ']');

内容的提问来源于stack exchange,提问作者ashrth

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 01:06:19