TensorFlow Hub加载ELMo模型触发UnicodeDecodeError,换版本无效求助
解决加载ELMo时的UnicodeDecodeError问题
问题重现
运行以下代码时触发编码解码错误:
import tensorflow.compat.v1 as tf tf.disable_v2_behavior() tf.disable_eager_execution() # Load pre trained ELMo model elmo = hub.Module("https://tfhub.dev/google/elmo/1", trainable=True)
报错信息:
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xc1 in position 149: invalid start byte
此前切换ELMo版本可解决该问题,但当前该方法已失效。
可行解决方案
清理TF Hub缓存
错误大概率由损坏的缓存模型文件引发,找到TF Hub默认缓存目录删除对应模型文件夹:- Linux/macOS路径:
~/.tfhub_modules/ - Windows路径:
C:\Users\<你的用户名>\.tfhub_modules\
删除google/elmo/1文件夹后重新运行代码,让TF Hub重新下载完整模型。
- Linux/macOS路径:
调整模型加载配置
修改加载代码,添加加载选项规避编码问题:import tensorflow.compat.v1 as tf import tensorflow_hub as hub tf.disable_v2_behavior() tf.disable_eager_execution() elmo = hub.Module( "https://tfhub.dev/google/elmo/1", trainable=True, load_options=tf.saved_model.LoadOptions(experimental_io_device="/job:localhost") )若已重新下载干净的本地模型,可直接指定本地路径加载:
elmo = hub.Module("/path/to/clean/elmo/1", trainable=True)回退兼容版本
尝试降级到此前能正常运行的TensorFlow和TF Hub版本组合,例如:pip install tensorflow==1.15 tensorflow-hub==0.12.0注意版本兼容性,TF Hub 0.12.0适配TensorFlow 1.x系列。
检查系统编码环境
在Jupyter中运行以下代码确认环境编码:import sys print(sys.getdefaultencoding()) print(sys.stdout.encoding)若默认编码非UTF-8,可通过设置环境变量
PYTHONIOENCODING=utf-8调整(Python 3中无法直接修改sys.setdefaultencoding)。
内容的提问来源于stack exchange,提问作者kwon_moment
相关产品推荐
相关产品推荐

