求助:将文本文件转为词频统计字典并找出高频词
词频统计代码问题修正
你的代码存在几处关键错误,导致无法正确统计词频:
- 变量名拼写错误:循环中使用了未定义的
last,实际应该是你创建的列表lst - 错误操作列表:试图用字符串作为索引给列表
lst赋值(lst[words] = ...),列表仅支持整数索引,不能这么用 dict.fromkeys用法完全错误:你传入的words是循环最后一个单词,freq始终为0,生成的字典和词频毫无关系- 文本处理不严谨:
split(" ")会在多空格场景下生成空字符串,且只处理了逗号和句号,其他标点未清理
修正方案1:手动实现词频统计
# 用with语句自动管理文件,避免资源泄漏 with open("Words.txt", "r") as file: word_list = [] for line in file: # 转小写并清理常见标点,可根据需求添加更多标点替换 cleaned_line = line.lower().replace(',', '').replace('.', '').replace('!', '').replace('?', '') # split()无参数时自动分割任意空白字符,同时忽略首尾空白 words = cleaned_line.split() word_list.extend(words) # 初始化词频字典 freq_dict = {} for word in word_list: if word in freq_dict: freq_dict[word] += 1 else: freq_dict[word] = 1 # 找出出现次数最多的单词 if freq_dict: most_common_word = max(freq_dict, key=freq_dict.get) print(f"出现次数最多的单词: {most_common_word},共出现{freq_dict[most_common_word]}次") else: print("文件中无有效单词") print("完整词频统计:", freq_dict)
修正方案2:用collections.Counter(更简洁高效)
Python标准库的Counter专门用于统计可哈希对象的出现次数,代码更简洁:
from collections import Counter import re with open("Words.txt", "r") as file: all_words = [] for line in file: # 用正则匹配所有单词,自动处理大小写和标点 words = re.findall(r'\b\w+\b', line.lower()) all_words.extend(words) freq_counter = Counter(all_words) # 获取出现次数最多的单词(most_common(1)返回包含元组的列表) if freq_counter: most_common_word, count = freq_counter.most_common(1)[0] print(f"出现次数最多的单词: {most_common_word},共出现{count}次") else: print("文件中无有效单词") print("完整词频统计:", freq_counter)
关键优化点说明
- 使用
with语句打开文件,无需手动调用close(),更安全 - 文本处理用
split()无参数或正则表达式,避免空字符串和遗漏标点的问题 - 用字典(或
Counter)专门存储词频,而非错误使用列表 - 用
max()配合key参数或Counter.most_common()快速找出高频词
内容的提问来源于stack exchange,提问作者SGYByeElo
相关产品推荐
相关产品推荐

