基于ELMo嵌入的CNN文本分类模型无参数输出问题求助
问题:ELMo+CNN文本分类模型参数显示为0的解决方法
尝试用推特数据结合ELMo嵌入构建CNN文本分类模型,模型未编译时打印摘要显示总参数、可训练参数、不可训练参数均为0,怀疑输入形状或填充问题,但无法推进。
原始实现代码
from sklearn.model_selection import train_test_split x = seeker_df[['tweet']] y = seeker_df['BinaryNumTarget'] #splitting into validation and test - 85% of dataset if train and validation, 15% for test x_train_val, x_test, y_train_val, y_test = train_test_split(x, y, test_size =0.15, stratify = y, # sampling used to ensure class distribution random_state = 42) #splitting into train and val - 70% for training, 15% for validation x_train, x_val, y_train, y_val = train_test_split(x_train_val, y_train_val, test_size =0.1765, stratify = y_train_val, random_state=42) #package installation from keras.models import Sequential from keras.layers import Conv1D, GlobalMaxPooling1D, Dense from keras.preprocessing.sequence import pad_sequences from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score from keras.callbacks import EarlyStopping #setting up parameters no_of_filters = 250 kernel_size = 3 hidden_dims = 250 batch_size = 32 epochs = 2 early_stopping = EarlyStopping(monitor ='val_loss', patience = 10) #obtaining elmo emmbedding elmo = hub.KerasLayer("https://tfhub.dev/google/elmo/2", trainable=True, name = 'elmo_embedding', input_shape=[], dtype=(tf.string)) embedding = tf.keras.layers.Lambda(elmo) model = Sequential() model.add(embedding) model.add(Conv1D(no_of_filters, kernel_size, activation ='relu')) model.add(GlobalMaxPooling1D()) model.add(Dense(hidden_dims, activation = 'relu')) model.add(Dense(1, activation = 'sigmoid')) model.summary()
当前模型摘要
Model: "sequential_5" ┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━┓ ┃ Layer (type) ┃ Output Shape ┃ Param # ┃ ┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━┩ │ lambda_3 (Lambda) │ ? │ 0 (unbuilt) │ ├──────────────────────────────────────┼─────────────────────────────┼─────────────────┤ │ conv1d_3 (Conv1D) │ ? │ 0 (unbuilt) │ ├──────────────────────────────────────┼─────────────────────────────┼─────────────────┤ │ global_max_pooling1d_3 │ ? │ 0 │ │ (GlobalMaxPooling1D) │ │ │ ├──────────────────────────────────────┼─────────────────────────────┼─────────────────┤ │ dense_6 (Dense) │ ? │ 0 (unbuilt) │ ├──────────────────────────────────────┼─────────────────────────────┼─────────────────┤ │ dense_7 (Dense) │ ? │ 0 (unbuilt) │ └──────────────────────────────────────┴─────────────────────────────┴─────────────────┘ Total params: 0 (0.00 B) Trainable params: 0 (0.00 B) Non-trainable params: 0 (0.00 B)
问题根源与解决方法
核心原因
Keras Sequential模型在未明确输入形状、未处理过真实数据时,无法自动推断各层的输出维度,导致参数无法计算,全部显示为0。你的代码存在三个关键问题:
- ELMo层用Lambda包裹,导致Keras无法识别该层的参数和输出形状;
- 输入数据是二维DataFrame列,不符合ELMo要求的一维字符串张量输入格式;
- ELMo输出是2D张量,而Conv1D需要3D张量输入,维度不匹配。
修正后的代码
import tensorflow as tf import tensorflow_hub as hub from sklearn.model_selection import train_test_split from keras.models import Sequential from keras.layers import Conv1D, GlobalMaxPooling1D, Dense from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score from keras.callbacks import EarlyStopping # 调整输入格式:从DataFrame列转为一维字符串数组 x = seeker_df['tweet'].values y = seeker_df['BinaryNumTarget'].values # 数据拆分逻辑不变 x_train_val, x_test, y_train_val, y_test = train_test_split(x, y, test_size=0.15, stratify=y, random_state=42) x_train, x_val, y_train, y_val = train_test_split(x_train_val, y_train_val, test_size=0.1765, stratify=y_train_val, random_state=42) # 参数设置不变 no_of_filters = 250 kernel_size = 3 hidden_dims = 250 batch_size = 32 epochs = 2 early_stopping = EarlyStopping(monitor='val_loss', patience=10) # 直接使用ELMo的KerasLayer,无需Lambda包裹 elmo_layer = hub.KerasLayer("https://tfhub.dev/google/elmo/2", trainable=True, name='elmo_embedding', input_shape=[], dtype=tf.string) model = Sequential() model.add(elmo_layer) # ELMo输出为(batch_size, 1024),Reshape为(batch_size, 1, 1024)适配Conv1D的3D输入要求 model.add(tf.keras.layers.Reshape((1, 1024))) model.add(Conv1D(no_of_filters, kernel_size, activation='relu')) model.add(GlobalMaxPooling1D()) model.add(Dense(hidden_dims, activation='relu')) model.add(Dense(1, activation='sigmoid')) # 现在可以正确显示参数 model.summary() # 编译并训练模型 model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy']) model.fit(x_train, y_train, batch_size=batch_size, epochs=epochs, validation_data=(x_val, y_val), callbacks=[early_stopping])
关键修正点说明
- 输入格式调整:将
x = seeker_df[['tweet']]改为x = seeker_df['tweet'].values,得到一维字符串数组,匹配ELMo的输入要求; - ELMo层正确使用:直接将
hub.KerasLayer加入模型,去掉Lambda包裹,让Keras能识别该层的参数和输出形状; - Reshape层适配:ELMo输出的是2D张量,Conv1D需要3D输入(
(batch_size, sequence_length, feature_dim)),通过Reshape将维度调整为(batch_size, 1, 1024),确保Conv1D可以正常处理。
内容的提问来源于stack exchange,提问作者leakie
相关产品推荐
相关产品推荐

