Python读取文件名触发UnicodeDecodeError,求解决方案
解决UnicodeDecodeError的方案
你的核心问题是读取文件时编码不匹配:Python默认用UTF-8解码文本文件,但你的部分文件可能使用了其他编码格式(比如Latin-1、GBK,或者系统默认的本地编码),导致遇到无法解码的字节0xe8。下面是具体的修复步骤:
1. 显式指定文件编码读取
最简单的临时修复是直接指定兼容所有字节的编码(比如latin-1),这样不会触发解码错误,之后你可以再排查具体编码:
# 替换原来的open语句 with open(file_path, encoding='latin-1') as f: first_line = f.readline()
如果latin-1显示的内容乱码,你可以用chardet库检测文件的真实编码:
# 先安装chardet:pip install chardet import chardet with open(file_path, 'rb') as f: result = chardet.detect(f.read()) # 用检测到的编码读取文件 with open(file_path, encoding=result['encoding']) as f: first_line = f.readline()
2. 修复代码中的其他小问题
你的代码还有几个可以优化的地方,虽然不是报错原因,但会影响运行效率和正确性:
- import语句不要放在循环里:
stopwords和SnowballStemmer的导入应该放在文件最开头,而不是for循环内部,不然每次循环都会重复导入模块。 - shuffle参数应该用布尔值:
sklearn.datasets.load_files的shuffle参数接受的是布尔值True/False,不是字符串'False',改成shuffle=False才会生效。
修复后的完整代码示例:
from __future__ import print_function import sklearn.datasets import nltk from os import listdir from os.path import isfile, join from nltk.corpus import stopwords from nltk.stem.snowball import SnowballStemmer # 初始化停用词和词干提取器 stopwords_en = stopwords.words('english') stemmer = SnowballStemmer("english") # 加载数据集,shuffle用布尔值 dataset = sklearn.datasets.load_files('data/', shuffle=False) categories = dataset.target_names for c in categories: directory_path = f'data/{c}' onlyfiles = [f for f in listdir(directory_path) if isfile(join(directory_path, f))] print(f"Level 1 Intent : {c}") print("---------------------------------------") for file_name in onlyfiles: file_path = join(directory_path, file_name) # 用chardet检测编码后读取 import chardet with open(file_path, 'rb') as f: detect_result = chardet.detect(f.read()) # 兼容检测失败的情况, fallback到latin-1 with open(file_path, encoding=detect_result['encoding'] or 'latin-1') as f: first_line = f.readline().strip() # 去掉换行符让输出更整洁 print(f"Level 2 Intent for {c} : {first_line}")
3. 额外排查建议
如果还是有问题,你可以检查一下报错的具体文件:
- 在循环中加入
try-except捕获错误,打印出触发问题的file_path,手动打开该文件确认是否有特殊字符或非UTF-8编码内容。 - 虽然你能正常输出文件名,但也可以确认下文件名是否包含非UTF-8字符(不过这种情况报错位置会在文件名读取阶段,而非文件内容)。
内容的提问来源于stack exchange,提问作者Nayantara Jeyaraj
相关产品推荐
相关产品推荐

