Keras如何加载带有预训练ELMO层的模型?两种加载方式均报错求解
问题原因分析
- 直接加载模型报错
ValueError: No model found in config file:因为模型中使用了封装ELMo的Lambda自定义层,Keras序列化保存整模型时无法自动存储自定义层的实现逻辑,加载时无法还原模型结构,因此整模型加载方案不适用。 - 手动构建结构后加载权重报错
AttributeError: 'str' object has no attribute 'decode':是h5py版本不兼容导致的。你使用的TensorFlow 1.15配套的Keras版本保存权重时,将元信息以字节格式存储在h5文件中,而3.x版本的h5py读取属性时会自动转为字符串,旧版Keras调用decode方法就会报错。 - 额外隐藏问题:你原代码中先加载权重,再执行全局变量初始化,会把已经加载的权重覆盖为随机值,就算没有版本错误也无法得到正确的预测结果。
解决方案
步骤1:调整依赖版本
执行命令降级h5py到兼容版本:
pip install h5py==2.10.0
步骤2:调整代码逻辑
调整session初始化和权重加载的顺序,加载权重时指定by_name=True保证层和权重对应,完整可运行代码如下:
import tensorflow_hub as hub import tensorflow as tf import numpy as np from tensorflow.keras.layers import Input, Lambda, Bidirectional, Dense, Dropout, LSTM from tensorflow.keras.models import Model # 定义ELMo嵌入逻辑 elmo = hub.Module("https://tfhub.dev/google/elmo/3", trainable=False) def ELMoEmbedding(input_text): return elmo(tf.reshape(tf.cast(input_text, tf.string), [-1]), signature="default", as_dict=True)["elmo"] def build_model(): input_layer = Input(shape=(1,), dtype="string", name="Input_layer") embedding_layer = Lambda(ELMoEmbedding, output_shape=(1024, ), name="Elmo_Embedding")(input_layer) BiLSTM = Bidirectional(LSTM(128, return_sequences= False, recurrent_dropout=0.2, dropout=0.2), name="BiLSTM")(embedding_layer) Dense_layer_1 = Dense(64, activation='relu')(BiLSTM) Dropout_layer_1 = Dropout(0.5)(Dense_layer_1) Dense_layer_2 = Dense(32, activation='relu')(Dropout_layer_1) Dropout_layer_2 = Dropout(0.5)(Dense_layer_2) output_layer = Dense(3, activation='sigmoid')(Dropout_layer_2) model = Model(inputs=[input_layer], outputs=output_layer, name="BiLSTM with ELMo Embeddings") model.compile(loss='binary_crossentropy',optimizer='adam',metrics=['accuracy']) return model # 先构建模型结构 elmo_BiDirectional_model = build_model() # 初始化会话,先执行初始化操作,再加载权重 with tf.Session() as session: session.run(tf.global_variables_initializer()) session.run(tf.tables_initializer()) # 加载权重,指定by_name=True避免层顺序不匹配的问题 elmo_BiDirectional_model.load_weights('weights.h5', by_name=True) # 测试预测 x_data = np.zeros((1, 1), dtype='object') x_data[0] = "test token" print(elmo_BiDirectional_model.predict(x_data))
内容的提问来源于stack exchange,提问作者webber
相关产品推荐
相关产品推荐

