Python新手求助:如何获取clean.txt中长度>3的高频词
解决获取长单词及其出现次数的问题
你的核心需求是筛选出长度大于3的单词并统计它们的出现次数,下面是简化且高效的实现方案:
优化后的代码
import re from collections import Counter def get_long_word_counts(): # 读取文件并提取所有单词(转为小写,避免大小写干扰统计) with open('clean.txt', 'r', encoding='utf-8') as f: text = f.read().lower() words = re.findall(r'\w+', text) # 过滤出长度大于3的单词 filtered_words = [word for word in words if len(word) > 3] # 统计单词出现次数,并按次数降序排列 word_counts = Counter(filtered_words).most_common() # 输出结果 for word, count in word_counts: print(f"{word}: {count}") # 调用函数 get_long_word_counts()
对原代码问题的说明
- 你原有的
read_data()函数只是逐个输出符合长度要求的单词,没有做次数统计,所以无法显示出现次数。 - 你已经用了
Counter(words).most_common(100),这个方法本身就会返回按出现次数降序排列的结果,后续的count.sort(key=sort_key, reverse=True)是多余操作。 - 用
re.findall(r'\w+', text)提取单词比split()更可靠,因为split()会把带标点的字符串当成一个单词(比如"hello,"会被完整保留),而正则能准确提取纯单词。
内容的提问来源于stack exchange,提问作者arketipi
相关产品推荐
相关产品推荐

