Python入门者求助:解决UnicodeDecodeError及词频排序问题
问题解决与代码优化
1. 解决UnicodeDecodeError错误
错误核心是文件编码并非默认的UTF-8,导致解码失败。两种可行解决方式:
- 尝试指定文件实际使用的编码(比如
latin-1、cp1252这类常见非UTF-8编码) - 若不确定编码,使用
errors='ignore'跳过无法解码的字节(会丢失少量字符,但保证程序正常运行)
2. 代码优化与词频排序实现
原代码存在三个问题:message.split缺少括号导致未执行分割、未处理单词大小写和标点(统计结果不准确)、未实现词频排序功能。以下是修正后的完整代码:
import os import string count = {} os.chdir('/Users/lritter/Desktop/Python') item = int(input('Which line would you like to evaluate? ')) print('You entered: ', item) # 处理编码问题:先尝试UTF-8,失败则用latin-1兼容更多字节 try: with open('Obama_speech.txt', encoding='utf-8') as file: lines = file.readlines() except UnicodeDecodeError: with open('Obama_speech.txt', encoding='latin-1') as file: lines = file.readlines() # 验证行号有效性,避免索引越界 if item < 0 or item >= len(lines): print(f"Error: Line number {item} is out of range. File has {len(lines)} lines.") else: message = lines[item] # 清洗单词:转小写、移除标点、分割成单词列表 translator = str.maketrans('', '', string.punctuation) cleaned_words = message.lower().translate(translator).split() # 统计长度≥5的单词出现次数 for word in cleaned_words: if len(word) >= 5: count[word] = count.get(word, 0) + 1 # 按出现次数从多到少排序 sorted_word_counts = sorted(count.items(), key=lambda x: x[1], reverse=True) print("\nWord counts (sorted by frequency descending):") for word, freq in sorted_word_counts: print(f"{word}: {freq}")
关键优化点说明
- 编码兼容:通过try-except自动适配编码,避免因编码未知导致的崩溃
- 输入校验:增加行号范围检查,防止用户输入无效行号引发索引错误
- 单词清洗:统一转小写、移除标点,确保"Hello"和"hello"、"world,"和"world"被视为同一个单词
- 排序实现:使用
sorted()函数,通过key=lambda x: x[1]指定按词频排序,reverse=True实现降序排列
内容的提问来源于stack exchange,提问作者Lauren Carole
相关产品推荐
相关产品推荐

