如何解决运行Python代码时出现的UnicodeEncodeError错误?
解决UnicodeEncodeError错误的方案
问题场景
运行以下Python代码时触发UnicodeEncodeError:
ins = insightsearch.Analysis( df='encodedreviews.csv', column_name="Text",vader=False ) # 生成可视化洞察的HTML文件 ins.review_analyze()
使用的CSV文件内容:
App_Name,rating,date,variation,Text,feedback Alexa,5,31-Jul-18,Charcoal Fabric ,Love my Echo!,1 Alexa,5,31-Jul-18,Charcoal Fabric ,Loved it!,1 Alexa,4,31-Jul-18,Walnut Finish ,"Sometimes while playing a game, you can answer a question correctly but Alexa says you got it wrong and answers the same as you. I like being able to turn lights on and off while away from home.",1 Alexa,5,31-Jul-18,Charcoal Fabric ,"I have had a lot of fun with this thing. My 4 yr old learns about dinosaurs, i control the lights and play games like categories. Has nice sound when playing music as well.",1 Alexa,5,31-Jul-18,Charcoal Fabric ,Music,1 Alexa,5,31-Jul-18,Heather Gray Fabric ,"I received the echo as a gift. I needed another Bluetooth or something to play music easily accessible, and found this smart speaker. Can’t wait to see what else it can do.",1
报错信息:
Exception has occurred: UnicodeEncodeError
'charmap' codec can't encode character '\u25c4' in position 282537: character maps to
File "C:\Users\rkono\OneDrive\Pictures\New folder\sen.py", line 12, in
ins.review_analyze()
UnicodeEncodeError: 'charmap' codec can't encode character '\u25c4' in position 282537: character maps to
解决方法
1. 手动以UTF-8编码读取CSV
insightsearch可能默认使用系统编码(Windows下多为cp1252)读取文件,无法处理部分Unicode字符。先手动读取CSV为DataFrame再传入:
import pandas as pd import insightsearch # 强制用UTF-8编码读取CSV df = pd.read_csv('encodedreviews.csv', encoding='utf-8') # 传入DataFrame而非文件名 ins = insightsearch.Analysis( df=df, column_name="Text", vader=False ) ins.review_analyze()
2. 设置系统输出编码为UTF-8
若错误出现在生成HTML的写入环节,在代码开头添加以下代码强制输出编码:
import sys sys.stdout.reconfigure(encoding='utf-8')
3. 清理文本中的特殊Unicode字符
如果上述方法无效,可过滤Text列中的特殊字符,保留常规可打印字符:
import pandas as pd import insightsearch import re df = pd.read_csv('encodedreviews.csv', encoding='utf-8') # 移除非文字、空格、常用标点外的特殊字符 df['Text'] = df['Text'].apply(lambda x: re.sub(r'[^\w\s.,!?\'"]', '', str(x))) ins = insightsearch.Analysis( df=df, column_name="Text", vader=False ) ins.review_analyze()
4. 检查库的编码参数
查看insightsearch的review_analyze()方法文档,若支持指定输出编码,直接传入参数:
# 若库支持该参数,尝试指定UTF-8编码输出 ins.review_analyze(output_encoding='utf-8')
内容的提问来源于stack exchange,提问作者NOORDEEN M
相关产品推荐
相关产品推荐

