使用TensorFlow Hub加载本地模型文件时遭遇UnicodeDecodeError错误的技术求助
解决TensorFlow Hub加载本地二进制文件时的UnicodeDecodeError问题
你遇到的这个UnicodeDecodeError本质原因很明确:你试图用TensorFlow Hub的hub.Module加载一个普通的预训练词向量文件,而不是TF Hub规范的模块文件。
先看你的代码:
import tensorflow_hub as hub module_url = "C:/study/web_mystem_skipgram_500_2_2015.bin.gz" embed = hub.Module(module_url)
这个web_mystem_skipgram_500_2_2015.bin.gz是Mystem发布的预训练词向量二进制文件,它并不符合TensorFlow Hub模块的格式要求——TF Hub模块需要是包含saved_model.pb和权重目录的SavedModel结构,而不是单一的二进制压缩文件。当hub.Module尝试读取这个文件时,会错误地把二进制内容当作UTF-8文本解析,从而触发解码失败。
解决方案分两种情况:
情况1:直接使用词向量,无需转成TF Hub模块
如果只是想在代码中使用这个词向量,推荐用gensim直接加载,再转换成TensorFlow可用的嵌入层:
import gzip from gensim.models import KeyedVectors import tensorflow as tf # 第一步:解压.gz文件 with gzip.open("C:/study/web_mystem_skipgram_500_2_2015.bin.gz", 'rb') as f_in: with open("web_mystem_skipgram_500_2_2015.bin", 'wb') as f_out: f_out.write(f_in.read()) # 第二步:用gensim加载词向量 word2vec_model = KeyedVectors.load_word2vec_format( "web_mystem_skipgram_500_2_2015.bin", binary=True ) # 第三步:转换成TensorFlow嵌入层(可直接用于模型) vocab_size = len(word2vec_model.index_to_key) embedding_dim = 500 embedding_layer = tf.keras.layers.Embedding( input_dim=vocab_size, output_dim=embedding_dim, weights=[word2vec_model.vectors], trainable=False # 如果不需要微调词向量,保持为False )
情况2:必须转成TF Hub模块使用
如果你的场景确实需要用TF Hub的方式加载,可以先把词向量封装成SavedModel格式,再用hub.Module加载:
import gzip from gensim.models import KeyedVectors import tensorflow as tf import tensorflow_hub as hub # 先加载词向量(同上步骤) with gzip.open("C:/study/web_mystem_skipgram_500_2_2015.bin.gz", 'rb') as f_in: with open("web_mystem_skipgram_500_2_2015.bin", 'wb') as f_out: f_out.write(f_in.read()) word2vec_model = KeyedVectors.load_word2vec_format( "web_mystem_skipgram_500_2_2015.bin", binary=True ) # 构建一个简单的TF模型,封装词向量 vocab = word2vec_model.index_to_key embedding_matrix = word2vec_model.vectors # 创建词汇表到索引的映射表 vocab_table = tf.lookup.StaticVocabularyTable( tf.lookup.KeyValueTensorInitializer( keys=vocab, values=tf.range(len(vocab), dtype=tf.int64) ), num_oov_buckets=1 # 处理未登录词 ) # 定义模型输入输出 input_text = tf.keras.layers.Input(shape=(), dtype=tf.string) text_indices = vocab_table.lookup(input_text) # 拼接未登录词的随机初始化向量(这里用0向量示例) full_embedding_matrix = tf.concat([embedding_matrix, tf.zeros((1, 500))], axis=0) embedding_output = tf.keras.layers.Embedding( input_dim=len(vocab)+1, output_dim=500, weights=[full_embedding_matrix], trainable=False )(text_indices) # 保存为SavedModel格式 hub_model = tf.keras.Model(inputs=input_text, outputs=embedding_output) hub_model.save("mystem_embedding_hub_module", save_format="tf") # 现在可以用hub.Module加载这个模块了 embed = hub.Module("mystem_embedding_hub_module")
额外提醒
以后使用hub.Module时,要确保加载的是TF Hub官方提供的模块URL,或者自己生成的符合SavedModel规范的本地目录,不要直接加载原始的词向量、权重文件这类非TF Hub格式的文件。
内容的提问来源于stack exchange,提问作者Andrew Alex
相关产品推荐
相关产品推荐

