You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取文件名触发UnicodeDecodeError,求解决方案

解决UnicodeDecodeError的方案

你的核心问题是读取文件时编码不匹配:Python默认用UTF-8解码文本文件,但你的部分文件可能使用了其他编码格式(比如Latin-1、GBK,或者系统默认的本地编码),导致遇到无法解码的字节0xe8。下面是具体的修复步骤:

1. 显式指定文件编码读取

最简单的临时修复是直接指定兼容所有字节的编码(比如latin-1),这样不会触发解码错误,之后你可以再排查具体编码:

# 替换原来的open语句
with open(file_path, encoding='latin-1') as f:
    first_line = f.readline()

如果latin-1显示的内容乱码,你可以用chardet库检测文件的真实编码:

# 先安装chardet:pip install chardet
import chardet

with open(file_path, 'rb') as f:
    result = chardet.detect(f.read())

# 用检测到的编码读取文件
with open(file_path, encoding=result['encoding']) as f:
    first_line = f.readline()

2. 修复代码中的其他小问题

你的代码还有几个可以优化的地方,虽然不是报错原因,但会影响运行效率和正确性:

  • import语句不要放在循环里:stopwords和SnowballStemmer的导入应该放在文件最开头,而不是for循环内部,不然每次循环都会重复导入模块。
  • shuffle参数应该用布尔值:sklearn.datasets.load_files的shuffle参数接受的是布尔值True/False,不是字符串'False',改成shuffle=False才会生效。

修复后的完整代码示例:

from __future__ import print_function
import sklearn.datasets
import nltk
from os import listdir
from os.path import isfile, join
from nltk.corpus import stopwords
from nltk.stem.snowball import SnowballStemmer

# 初始化停用词和词干提取器
stopwords_en = stopwords.words('english')
stemmer = SnowballStemmer("english")

# 加载数据集,shuffle用布尔值
dataset = sklearn.datasets.load_files('data/', shuffle=False)
categories = dataset.target_names

for c in categories:
    directory_path = f'data/{c}'
    onlyfiles = [f for f in listdir(directory_path) if isfile(join(directory_path, f))]
    print(f"Level 1 Intent : {c}")
    print("---------------------------------------")
    for file_name in onlyfiles:
        file_path = join(directory_path, file_name)
        # 用chardet检测编码后读取
        import chardet
        with open(file_path, 'rb') as f:
            detect_result = chardet.detect(f.read())
        # 兼容检测失败的情况, fallback到latin-1
        with open(file_path, encoding=detect_result['encoding'] or 'latin-1') as f:
            first_line = f.readline().strip()  # 去掉换行符让输出更整洁
        print(f"Level 2 Intent for {c} : {first_line}")

3. 额外排查建议

如果还是有问题,你可以检查一下报错的具体文件:

  • 在循环中加入try-except捕获错误,打印出触发问题的file_path,手动打开该文件确认是否有特殊字符或非UTF-8编码内容。
  • 虽然你能正常输出文件名,但也可以确认下文件名是否包含非UTF-8字符(不过这种情况报错位置会在文件名读取阶段,而非文件内容)。

内容的提问来源于stack exchange,提问作者Nayantara Jeyaraj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:56:58