You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow 2.14保存模型时出现UnicodeEncodeError编码错误

解决TensorFlow模型保存时的UnicodeEncodeError问题

问题背景

使用Python 3.9 + TensorFlow 2.14运行官方文本分类教程,模型训练和预测均正常,但调用export_model.save("exported.keras")保存模型时触发如下错误:

UnicodeEncodeError: 'charmap' codec can't encode character '\x96' in position 3325: character maps to <undefined>

错误栈指向模型层保存词汇表的代码段,崩溃发生在open文件后执行f.write的步骤:

def save_assets(self, dir_path):
    if self.input_vocabulary:
        return
    vocabulary = self.get_vocabulary(include_special_tokens=True)
    vocabulary_filepath = tf.io.gfile.join(dir_path, "vocabulary.txt")
    with open(vocabulary_filepath, "w") as f:
        f.write("\n".join([str(w) for w in vocabulary]))

解决方案

1. 定位引发错误的字符

在保存模型前,遍历词汇表找出无法用系统默认编码(Windows下为cp1252)编码的字符:

# 获取模型的词汇表(假设文本预处理层是模型第一层)
vocabulary = export_model.layers[0].get_vocabulary(include_special_tokens=True)

# 检查每个词汇的编码兼容性
for idx, word in enumerate(vocabulary):
    try:
        # 模拟默认编码尝试
        word.encode('cp1252')
    except UnicodeEncodeError:
        print(f"无法编码的条目索引: {idx}, 内容: {repr(word)}")

运行后即可定位到包含\x96(长破折号)的具体词汇。

2. 修改文件编码解决保存问题

问题根源是Windows系统默认编码不支持部分Unicode字符,指定utf-8编码写入文件即可修复:

方式一:修改TensorFlow源码(直接有效)

找到TensorFlow安装目录中对应层的save_assets方法(路径参考报错信息中的keras\src相关文件),将open调用修改为:

with open(vocabulary_filepath, "w", encoding='utf-8') as f:
    f.write("\n".join([str(w) for w in vocabulary]))

方式二:通过环境变量强制默认编码

在代码开头添加以下代码,无需修改源码即可强制Python使用utf-8作为默认编码:

import sys
import os

sys.stdout.reconfigure(encoding='utf-8')
os.environ['PYTHONUTF8'] = '1'

内容的提问来源于stack exchange,提问作者10mjg

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 08:10:13