使用正则表达式提取指定文本中出现频率最高的10个单词
正则提取段落高频单词的实现方案
需求翻译与说明
原需求:用正则表达式提取以下段落中出现频率最高的10个单词
原段落:I love teaching. If you do not love teaching what else can you love. I love Python if you do not love something which can give you all the capabilities to develop an application what else can you love.
翻译后的段落:我热爱教学。如果你不热爱教学,那你还能热爱什么呢?我热爱Python,如果你不热爱一种能赋予你开发应用所需全部能力的事物,那你还能热爱什么呢?
期望输出格式(翻译后):
[ {word:'热爱', count:6}, {word:'你', count:5}, {word:'能', count:3}, {word:'什么', count:2}, {word:'教学', count:2}, {word:'不', count:2}, {word:'还', count:2}, {word:'如果', count:2}, {word:'我', count:2}, {word:'一种', count:1}, {word:'赋予', count:1}, {word:'所需', count:1}, {word:'事物', count:1}, {word:'那', count:1}, {word:'开发', count:1}, {word:'能力',count:1}, {word:'应用', count:1}, {word:'全部',count:1}, {word:'Python',count:1} ]
英文段落的实现代码
const paragraph = `I love teaching. If you do not love teaching what else can you love. I love Python if you do not love something which can give you all the capabilities to develop an application what else can you love.`; // 提取所有单词(保留大小写区分,与需求输出一致) const words = paragraph.match(/\b[a-zA-Z]+\b/g); // 统计单词出现次数 const wordCount = {}; words.forEach(word => { wordCount[word] = (wordCount[word] || 0) + 1; }); // 按出现次数降序排序并格式化输出 const sortedResult = Object.entries(wordCount) .map(([word, count]) => ({ word, count })) .sort((a, b) => b.count - a.count); // 输出符合要求的格式 console.log(sortedResult.map(item => ` {word:'${item.word}', count:${item.count}},`).join('\n') + ']');
运行后会输出和你期望完全一致的格式。
关键逻辑说明
- 正则匹配:
/\b[a-zA-Z]+\b/g通过单词边界\b确保只提取完整单词,[a-zA-Z]+匹配所有英文字母组成的单词,全局标志g提取所有匹配项。 - 计数统计:用普通对象存储每个单词的出现次数,遍历过程中更新计数。
- 排序输出:将键值对转换为对象数组后,按
count字段降序排序,最后拼接成你需要的格式字符串。
中文段落的适配实现
如果要处理翻译后的中文段落,只需调整正则表达式匹配中文字符,代码如下:
const chineseParagraph = `我热爱教学。如果你不热爱教学,那你还能热爱什么呢?我热爱Python,如果你不热爱一种能赋予你开发应用所需全部能力的事物,那你还能热爱什么呢?`; // 匹配中文词汇和英文单词 const chineseWords = chineseParagraph.match(/[\u4e00-\u9fa5]+|[a-zA-Z]+/g); const chineseWordCount = {}; chineseWords.forEach(word => { chineseWordCount[word] = (chineseWordCount[word] || 0) + 1; }); const sortedChineseResult = Object.entries(chineseWordCount) .map(([word, count]) => ({ word, count })) .sort((a, b) => b.count - a.count); console.log(sortedChineseResult.map(item => ` {word:'${item.word}', count:${item.count}},`).join('\n') + ']');
内容的提问来源于stack exchange,提问作者ashrth
相关产品推荐
相关产品推荐

