You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

含中文数据的随机森林Graphviz可视化Unicode编码错误求助

解决随机森林决策树可视化时的UnicodeEncodeError(中文字符问题)

我基于含中文字符的PC订单数据训练了随机森林模型,建模与精度验证已完成,但生成决策树可视化图像时触发UnicodeEncodeError,推测是数据集的中文字符特征名导致。尝试过StringIO和BytesIO均无效,相关代码及报错信息如下:

导入代码

from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier
from sklearn.metrics import accuracy_score
from sklearn.preprocessing import OneHotEncoder
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn import tree
from sklearn.tree import export_graphviz
import pydot
from IPython.display import Image
import six
import sys
sys.modules['sklearn.externals.six'] = six
from io import StringIO,BytesIO

随机森林建模代码

from sklearn.ensemble import RandomForestClassifier

X = finaldata.drop(columns=['是否赢单'])  
y = finaldata['是否赢单']

categorical_cols = X.select_dtypes(include=['object']).columns
numerical_cols = X.select_dtypes(include=['number']).columns

preprocessor = ColumnTransformer(
    transformers=[
        ('num', 'passthrough', numerical_cols),
        ('cat', OneHotEncoder(handle_unknown='ignore'), categorical_cols)
    ])

clf = Pipeline(steps=[
    ('preprocessor', preprocessor),
    ('classifier', RandomForestClassifier(random_state=42))
])

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)

clf.fit(X_train, y_train)

y_pred = clf.predict(X_test)

绘图代码

feature_names = clf.named_steps['preprocessor'].get_feature_names_out()

single_tree = clf.named_steps['classifier'].estimators_[0]

dot_data = StringIO()
export_graphviz(single_tree, out_file=dot_data, 
                filled=True, rounded=True,
                special_characters=True,
                feature_names=feature_names, 
                class_names=['Loss', 'Win'])

dot_data_str = dot_data.getvalue()

(graph,) = pydot.graph_from_dot_data(dot_data_str)
graph.write_png('decision_tree.png')

Image(filename='decision_tree.png')

报错信息

UnicodeEncodeError                        Traceback (most recent call last)
Cell In [18], line 13
     11 # Draw the graph using pydot
     12 (graph,) = pydot.graph_from_dot_data(dot_data.getvalue())
---> 13 graph.write_png('decision_tree.png')
     15 # Display the image
     16 Image(filename='decision_tree.png')

File c:\Users\Theodore\AppData\Local\Programs\Python\Python310\lib\site-packages\pydot.py:1743, in Dot.__init__.<locals>.new_method(path, f, prog, encoding)
   1739 def new_method(
   1740         path, f=frmt, prog=self.prog,
   1741         encoding=None):
   1742     """Refer to docstring of method `write.`"""
-> 1743     self.write(
   1744         path, format=f, prog=prog,
   1745         encoding=encoding)

File c:\Users\Theodore\AppData\Local\Programs\Python\Python310\lib\site-packages\pydot.py:1828, in Dot.write(self, path, prog, format, encoding)
   1826         f.write(s)
   1827 else:
-> 1828     s = self.create(prog, format, encoding=encoding)
   1829     with io.open(path, mode='wb') as f:
   1830         f.write(s)
...
File c:\Users\Theodore\AppData\Local\Programs\Python\Python310\lib\encodings\cp1252.py:19, in IncrementalEncoder.encode(self, input, final)
     18 def encode(self, input, final=False):
---> 19     return codecs.charmap_encode(input,self.errors,encoding_table)[0]

UnicodeEncodeError: 'charmap' codec can't encode characters in position 163-166: character maps to <undefined>

解决方法

方法1:给pydot写入方法指定UTF-8编码

问题根源是pydot默认使用系统编码(如Windows的cp1252)无法处理中文字符,显式指定encoding='utf-8'即可解决:

feature_names = clf.named_steps['preprocessor'].get_feature_names_out()

single_tree = clf.named_steps['classifier'].estimators_[0]

dot_data = StringIO()
export_graphviz(single_tree, out_file=dot_data, 
                filled=True, rounded=True,
                special_characters=True,
                feature_names=feature_names, 
                class_names=['Loss', 'Win'])

dot_data_str = dot_data.getvalue()

(graph,) = pydot.graph_from_dot_data(dot_data_str)
# 关键:添加encoding='utf-8'参数
graph.write_png('decision_tree.png', encoding='utf-8')

Image(filename='decision_tree.png')

方法2:改用sklearn内置plot_tree绘图(更稳定支持中文)

绕过pydot,直接使用sklearn自带的tree.plot_tree,只需配置中文字体即可:

import matplotlib.pyplot as plt
from sklearn import tree

# 配置中文字体,替换为你系统支持的字体(如Mac用'Arial Unicode MS')
plt.rcParams['font.sans-serif'] = ['SimHei']
plt.rcParams['axes.unicode_minus'] = False

feature_names = clf.named_steps['preprocessor'].get_feature_names_out()
single_tree = clf.named_steps['classifier'].estimators_[0]

# 设置画布大小
plt.figure(figsize=(25, 12))
tree.plot_tree(single_tree,
               filled=True, rounded=True,
               feature_names=feature_names,
               class_names=['输单', '赢单'])  # 可以改成中文类名
# 保存图像,bbox_inches='tight'防止文字被截断
plt.savefig('decision_tree.png', dpi=300, bbox_inches='tight')
plt.show()

方法3:在导出dot数据时指定UTF-8编码

通过BytesIO存储dot数据,导出时显式指定编码,确保字符串正确解析:

feature_names = clf.named_steps['preprocessor'].get_feature_names_out()

single_tree = clf.named_steps['classifier'].estimators_[0]

# 使用BytesIO并指定utf-8编码
dot_data = BytesIO()
export_graphviz(single_tree, out_file=dot_data, 
                filled=True, rounded=True,
                special_characters=True,
                feature_names=feature_names, 
                class_names=['Loss', 'Win'],
                encoding='utf-8')  # 关键:指定编码

# 解码成utf-8字符串
dot_data_str = dot_data.getvalue().decode('utf-8')

(graph,) = pydot.graph_from_dot_data(dot_data_str)
graph.write_png('decision_tree.png', encoding='utf-8')

Image(filename='decision_tree.png')

内容的提问来源于stack exchange,提问作者Theodore Maximus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 01:32:04