You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python入门者求助:解决UnicodeDecodeError及词频排序问题

问题解决与代码优化

1. 解决UnicodeDecodeError错误

错误核心是文件编码并非默认的UTF-8,导致解码失败。两种可行解决方式:

  • 尝试指定文件实际使用的编码(比如latin-1、cp1252这类常见非UTF-8编码)
  • 若不确定编码,使用errors='ignore'跳过无法解码的字节(会丢失少量字符,但保证程序正常运行)

2. 代码优化与词频排序实现

原代码存在三个问题:message.split缺少括号导致未执行分割、未处理单词大小写和标点(统计结果不准确)、未实现词频排序功能。以下是修正后的完整代码:

import os
import string

count = {}
os.chdir('/Users/lritter/Desktop/Python')

item = int(input('Which line would you like to evaluate? '))
print('You entered: ', item)

# 处理编码问题:先尝试UTF-8,失败则用latin-1兼容更多字节
try:
    with open('Obama_speech.txt', encoding='utf-8') as file:
        lines = file.readlines()
except UnicodeDecodeError:
    with open('Obama_speech.txt', encoding='latin-1') as file:
        lines = file.readlines()

# 验证行号有效性,避免索引越界
if item < 0 or item >= len(lines):
    print(f"Error: Line number {item} is out of range. File has {len(lines)} lines.")
else:
    message = lines[item]
    # 清洗单词:转小写、移除标点、分割成单词列表
    translator = str.maketrans('', '', string.punctuation)
    cleaned_words = message.lower().translate(translator).split()

    # 统计长度≥5的单词出现次数
    for word in cleaned_words:
        if len(word) >= 5:
            count[word] = count.get(word, 0) + 1

    # 按出现次数从多到少排序
    sorted_word_counts = sorted(count.items(), key=lambda x: x[1], reverse=True)

    print("\nWord counts (sorted by frequency descending):")
    for word, freq in sorted_word_counts:
        print(f"{word}: {freq}")

关键优化点说明

  • 编码兼容:通过try-except自动适配编码,避免因编码未知导致的崩溃
  • 输入校验:增加行号范围检查,防止用户输入无效行号引发索引错误
  • 单词清洗:统一转小写、移除标点,确保"Hello"和"hello"、"world,"和"world"被视为同一个单词
  • 排序实现:使用sorted()函数,通过key=lambda x: x[1]指定按词频排序,reverse=True实现降序排列

内容的提问来源于stack exchange,提问作者Lauren Carole

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 03:20:14