如何使用Python统计文本中各单词数量?附失效代码求助
原代码问题点
原代码无法正常统计的核心原因有两个:
- 直接用
for word in line遍历文本行时,实际是逐字符遍历字符串,永远匹配不到完整的"Python"单词,计数结果始终为0 - 逻辑仅针对单一字符串"Python"做判断,没有实现全量单词的计数能力,也没有处理单词附带标点、大小写不一致导致的匹配误差问题
修正代码
版本1:无额外依赖,统计所有单词出现次数
import string file_path = "file.txt" word_count = {} with open(file_path, 'r', encoding='utf-8') as file_object: for line in file_object: # 统一转小写,消除大小写差异带来的计数错误 line = line.lower() # 按空白字符拆分得到单词列表 for word in line.split(): # 清除单词首尾附着的标点,比如"python,"、"python."这类情况 clean_word = word.strip(string.punctuation) if not clean_word: continue # 累加计数 word_count[clean_word] = word_count.get(clean_word, 0) + 1 # 输出结果 for word, cnt in word_count.items(): print(f"{word}: {cnt}")
版本2:仅统计特定单词(比如原代码要找的Python)的出现次数
import string file_path = "file.txt" target = "python" cnt = 0 with open(file_path, 'r', encoding='utf-8') as file_object: for line in file_object: line = line.lower() for word in line.split(): if word.strip(string.punctuation) == target: cnt += 1 print(f"'{target}' 共出现{cnt}次")
版本3:用标准库Counter简化代码
如果不需要避免额外引入标准库,可以直接用collections.Counter做计数,代码更简洁,还支持快速查询出现频率最高的单词:
import string from collections import Counter file_path = "file.txt" words = [] with open(file_path, 'r', encoding='utf-8') as file_object: for line in file_object: line = line.lower() for word in line.split(): clean_word = word.strip(string.punctuation) if clean_word: words.append(clean_word) word_count = Counter(words) # 输出全量计数 print(word_count) # 比如要查出现次数最高的3个单词,直接调用most_common即可 # print(word_count.most_common(3))
关键注意点
- 遍历字符串对象时,迭代返回的是单个字符,要拿到单词必须先用
split()方法按空白拆分 - 文本统计时建议统一做大小写转换、标点清理,否则同个单词的不同形态会被识别为不同内容,导致计数不准
- 打开文件时建议显式指定文件编码,避免不同系统默认编码不一致导致的读取报错
内容的提问来源于stack exchange,提问作者Dejo chalkiadaky
相关产品推荐
相关产品推荐

