You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

运行情感词典Python代码报'utf-8' codec can't decode byte 0xf3错误如何解决

报错根因

该UnicodeDecodeError由文件编码不匹配导致:读取positive.txt、negative.txt时Python默认使用utf-8编码解码,但你的两个语料文件实际编码并非utf-8,字节0xf3不属于utf-8合法字符范围,因此触发解码失败。

修复方法

方法一:指定兼容编码打开(推荐)

新增encoding='latin-1'参数打开文件即可,latin-1为单字节编码,可兼容所有单字节字符,不会出现解码报错。同时可直接遍历文件对象逐行读取,避免大文件读取时内存占用过高:

from textblob import TextBlob

pos_count = 0
pos_correct = 0

with open("positive.txt","r", encoding='latin-1') as f:
    for line in f:
        analysis = TextBlob(line)
        if analysis.sentiment.polarity > 0:
            pos_correct += 1
        pos_count +=1


neg_count = 0
neg_correct = 0

with open("negative.txt","r", encoding='latin-1') as f:
    for line in f:
        analysis = TextBlob(line)
        if analysis.sentiment.polarity <= 0:
            neg_correct += 1
        neg_count +=1

print("Positive accuracy = {}% via {} samples".format(pos_correct/pos_count*100.0, pos_count))
print("Negative accuracy = {}% via {} samples".format(neg_correct/neg_count*100.0, neg_count))

方法二:忽略解码错误

如果确认部分异常字符不影响最终准确率,也可添加errors='ignore'参数,直接跳过无法解码的字节:

# 示例写法,两个文件打开都可添加该参数
with open("positive.txt","r", encoding='utf-8', errors='ignore') as f:

内容的提问来源于stack exchange,提问作者Radhika Singh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 16:24:04