Python读取文本文件统计词频并按字母序排序输出字典实现问题
Python词频统计实现方案
完整实现代码(基于标准库,简洁高效)
from collections import Counter def movietext(): with open("movietext.txt", "r") as textfile: text = textfile.read() # 清洗文本:移除感叹号、句号等特殊字符 for char in "!.": text = text.replace(char, "") return text if __name__ == "__main__": # 读取并清洗文本 cleaned_text = movietext() # 分割为单词列表 word_list = cleaned_text.split() # 统计单词出现频次 word_count = Counter(word_list) # 按单词字母顺序排序,生成保序字典 sorted_word_dict = dict(sorted(word_count.items())) # 按要求格式输出 for word, count in sorted_word_dict.items(): print(f"{word}: {count}")
无需额外导入的手动实现版本
def movietext(): with open("movietext.txt", "r") as textfile: text = textfile.read() for char in "!.": text = text.replace(char, "") return text if __name__ == "__main__": cleaned_text = movietext() word_list = cleaned_text.split() word_count = {} # 手动统计词频 for word in word_list: word_count[word] = word_count.get(word, 0) + 1 # 按字母顺序排序后输出 for word, count in sorted(word_count.items()): print(f"{word}: {count}")
关键说明
- 文本清洗环节补充了对句号的替换,避免原文本中
go.的句号被计入单词 split()默认按任意空白字符(空格、换行符等)分割文本,刚好匹配单词拆分需求- 默认排序规则完全符合要求的字母顺序,大写单词会排在同首字母小写单词前,和预期输出格式一致
- Python 3.7及以上版本中,
dict(sorted(...))生成的字典会严格保留排序顺序,符合输出字典的要求
内容的提问来源于stack exchange,提问作者Dominik
相关产品推荐
相关产品推荐

