运行情感词典Python代码报'utf-8' codec can't decode byte 0xf3错误如何解决
报错根因
该UnicodeDecodeError由文件编码不匹配导致:读取positive.txt、negative.txt时Python默认使用utf-8编码解码,但你的两个语料文件实际编码并非utf-8,字节0xf3不属于utf-8合法字符范围,因此触发解码失败。
修复方法
方法一:指定兼容编码打开(推荐)
新增encoding='latin-1'参数打开文件即可,latin-1为单字节编码,可兼容所有单字节字符,不会出现解码报错。同时可直接遍历文件对象逐行读取,避免大文件读取时内存占用过高:
from textblob import TextBlob pos_count = 0 pos_correct = 0 with open("positive.txt","r", encoding='latin-1') as f: for line in f: analysis = TextBlob(line) if analysis.sentiment.polarity > 0: pos_correct += 1 pos_count +=1 neg_count = 0 neg_correct = 0 with open("negative.txt","r", encoding='latin-1') as f: for line in f: analysis = TextBlob(line) if analysis.sentiment.polarity <= 0: neg_correct += 1 neg_count +=1 print("Positive accuracy = {}% via {} samples".format(pos_correct/pos_count*100.0, pos_count)) print("Negative accuracy = {}% via {} samples".format(neg_correct/neg_count*100.0, neg_count))
方法二:忽略解码错误
如果确认部分异常字符不影响最终准确率,也可添加errors='ignore'参数,直接跳过无法解码的字节:
# 示例写法,两个文件打开都可添加该参数 with open("positive.txt","r", encoding='utf-8', errors='ignore') as f:
内容的提问来源于stack exchange,提问作者Radhika Singh
相关产品推荐
相关产品推荐

