You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决运行Python代码时出现的UnicodeEncodeError错误?

解决UnicodeEncodeError错误的方案

问题场景

运行以下Python代码时触发UnicodeEncodeError:

ins = insightsearch.Analysis(
  df='encodedreviews.csv',
  column_name="Text",vader=False
)

# 生成可视化洞察的HTML文件
ins.review_analyze()

使用的CSV文件内容:

App_Name,rating,date,variation,Text,feedback
Alexa,5,31-Jul-18,Charcoal Fabric ,Love my Echo!,1
Alexa,5,31-Jul-18,Charcoal Fabric ,Loved it!,1
Alexa,4,31-Jul-18,Walnut Finish ,"Sometimes while playing a game, you can answer a question correctly but Alexa says you got it wrong and answers the same as you.  I like being able to turn lights on and off while away from home.",1
Alexa,5,31-Jul-18,Charcoal Fabric ,"I have had a lot of fun with this thing. My 4 yr old learns about dinosaurs, i control the lights and play games like categories. Has nice sound when playing music as well.",1
Alexa,5,31-Jul-18,Charcoal Fabric ,Music,1
Alexa,5,31-Jul-18,Heather Gray Fabric ,"I received the echo as a gift. I needed another Bluetooth or something to play music easily accessible, and found this smart speaker. Can’t wait to see what else it can do.",1

报错信息:

Exception has occurred: UnicodeEncodeError
'charmap' codec can't encode character '\u25c4' in position 282537: character maps to
File "C:\Users\rkono\OneDrive\Pictures\New folder\sen.py", line 12, in
ins.review_analyze()
UnicodeEncodeError: 'charmap' codec can't encode character '\u25c4' in position 282537: character maps to

解决方法

1. 手动以UTF-8编码读取CSV

insightsearch可能默认使用系统编码(Windows下多为cp1252)读取文件,无法处理部分Unicode字符。先手动读取CSV为DataFrame再传入:

import pandas as pd
import insightsearch

# 强制用UTF-8编码读取CSV
df = pd.read_csv('encodedreviews.csv', encoding='utf-8')

# 传入DataFrame而非文件名
ins = insightsearch.Analysis(
  df=df,
  column_name="Text",
  vader=False
)

ins.review_analyze()

2. 设置系统输出编码为UTF-8

若错误出现在生成HTML的写入环节,在代码开头添加以下代码强制输出编码:

import sys
sys.stdout.reconfigure(encoding='utf-8')

3. 清理文本中的特殊Unicode字符

如果上述方法无效,可过滤Text列中的特殊字符,保留常规可打印字符:

import pandas as pd
import insightsearch
import re

df = pd.read_csv('encodedreviews.csv', encoding='utf-8')

# 移除非文字、空格、常用标点外的特殊字符
df['Text'] = df['Text'].apply(lambda x: re.sub(r'[^\w\s.,!?\'"]', '', str(x)))

ins = insightsearch.Analysis(
  df=df,
  column_name="Text",
  vader=False
)

ins.review_analyze()

4. 检查库的编码参数

查看insightsearch的review_analyze()方法文档,若支持指定输出编码,直接传入参数:

# 若库支持该参数,尝试指定UTF-8编码输出
ins.review_analyze(output_encoding='utf-8')

内容的提问来源于stack exchange,提问作者NOORDEEN M

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 19:04:55