You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于ELMo嵌入的CNN文本分类模型无参数输出问题求助

问题:ELMo+CNN文本分类模型参数显示为0的解决方法

尝试用推特数据结合ELMo嵌入构建CNN文本分类模型,模型未编译时打印摘要显示总参数、可训练参数、不可训练参数均为0,怀疑输入形状或填充问题,但无法推进。

原始实现代码

from sklearn.model_selection import train_test_split

x = seeker_df[['tweet']]
y = seeker_df['BinaryNumTarget']

#splitting into validation and test - 85% of dataset if train and validation, 15% for test
x_train_val, x_test, y_train_val, y_test = train_test_split(x,
                                                            y, 
                                                            test_size =0.15,
                                                            stratify = y, # sampling used to ensure class distribution
                                                            random_state = 42)

#splitting into train and val - 70% for training, 15% for validation
x_train, x_val, y_train, y_val = train_test_split(x_train_val, 
                                                  y_train_val, 
                                                  test_size =0.1765,
                                                  stratify = y_train_val, 
                                                  random_state=42)
#package installation 
from keras.models import Sequential 
from keras.layers import Conv1D, GlobalMaxPooling1D, Dense
from keras.preprocessing.sequence import pad_sequences
from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score
from keras.callbacks import EarlyStopping

#setting up parameters
no_of_filters = 250
kernel_size = 3
hidden_dims = 250
batch_size = 32
epochs = 2
early_stopping = EarlyStopping(monitor ='val_loss', patience = 10)

#obtaining elmo emmbedding
elmo = hub.KerasLayer("https://tfhub.dev/google/elmo/2", 
                      trainable=True,
                      name = 'elmo_embedding',
                      input_shape=[],
                     dtype=(tf.string))
embedding = tf.keras.layers.Lambda(elmo)

model = Sequential()

model.add(embedding)

model.add(Conv1D(no_of_filters, kernel_size, activation ='relu'))

model.add(GlobalMaxPooling1D())

model.add(Dense(hidden_dims, activation = 'relu'))

model.add(Dense(1, activation = 'sigmoid'))

model.summary()

当前模型摘要

Model: "sequential_5"
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━┓
┃ Layer (type)                         ┃ Output Shape                ┃         Param # ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━┩
│ lambda_3 (Lambda)                    │ ?                           │     0 (unbuilt) │
├──────────────────────────────────────┼─────────────────────────────┼─────────────────┤
│ conv1d_3 (Conv1D)                    │ ?                           │     0 (unbuilt) │
├──────────────────────────────────────┼─────────────────────────────┼─────────────────┤
│ global_max_pooling1d_3               │ ?                           │               0 │
│ (GlobalMaxPooling1D)                 │                             │                 │
├──────────────────────────────────────┼─────────────────────────────┼─────────────────┤
│ dense_6 (Dense)                      │ ?                           │     0 (unbuilt) │
├──────────────────────────────────────┼─────────────────────────────┼─────────────────┤
│ dense_7 (Dense)                      │ ?                           │     0 (unbuilt) │
└──────────────────────────────────────┴─────────────────────────────┴─────────────────┘
 Total params: 0 (0.00 B)
 Trainable params: 0 (0.00 B)
 Non-trainable params: 0 (0.00 B)

问题根源与解决方法

核心原因

Keras Sequential模型在未明确输入形状、未处理过真实数据时,无法自动推断各层的输出维度,导致参数无法计算,全部显示为0。你的代码存在三个关键问题:

  1. ELMo层用Lambda包裹,导致Keras无法识别该层的参数和输出形状;
  2. 输入数据是二维DataFrame列,不符合ELMo要求的一维字符串张量输入格式;
  3. ELMo输出是2D张量,而Conv1D需要3D张量输入,维度不匹配。

修正后的代码

import tensorflow as tf
import tensorflow_hub as hub
from sklearn.model_selection import train_test_split
from keras.models import Sequential 
from keras.layers import Conv1D, GlobalMaxPooling1D, Dense
from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score
from keras.callbacks import EarlyStopping

# 调整输入格式:从DataFrame列转为一维字符串数组
x = seeker_df['tweet'].values
y = seeker_df['BinaryNumTarget'].values

# 数据拆分逻辑不变
x_train_val, x_test, y_train_val, y_test = train_test_split(x,
                                                            y, 
                                                            test_size=0.15,
                                                            stratify=y,
                                                            random_state=42)

x_train, x_val, y_train, y_val = train_test_split(x_train_val, 
                                                  y_train_val, 
                                                  test_size=0.1765,
                                                  stratify=y_train_val, 
                                                  random_state=42)

# 参数设置不变
no_of_filters = 250
kernel_size = 3
hidden_dims = 250
batch_size = 32
epochs = 2
early_stopping = EarlyStopping(monitor='val_loss', patience=10)

# 直接使用ELMo的KerasLayer,无需Lambda包裹
elmo_layer = hub.KerasLayer("https://tfhub.dev/google/elmo/2", 
                            trainable=True,
                            name='elmo_embedding',
                            input_shape=[],
                            dtype=tf.string)

model = Sequential()
model.add(elmo_layer)
# ELMo输出为(batch_size, 1024),Reshape为(batch_size, 1, 1024)适配Conv1D的3D输入要求
model.add(tf.keras.layers.Reshape((1, 1024)))
model.add(Conv1D(no_of_filters, kernel_size, activation='relu'))
model.add(GlobalMaxPooling1D())
model.add(Dense(hidden_dims, activation='relu'))
model.add(Dense(1, activation='sigmoid'))

# 现在可以正确显示参数
model.summary()

# 编译并训练模型
model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])
model.fit(x_train, y_train, 
          batch_size=batch_size, 
          epochs=epochs, 
          validation_data=(x_val, y_val),
          callbacks=[early_stopping])

关键修正点说明

  • 输入格式调整:将x = seeker_df[['tweet']]改为x = seeker_df['tweet'].values,得到一维字符串数组,匹配ELMo的输入要求;
  • ELMo层正确使用:直接将hub.KerasLayer加入模型,去掉Lambda包裹,让Keras能识别该层的参数和输出形状;
  • Reshape层适配:ELMo输出的是2D张量,Conv1D需要3D输入((batch_size, sequence_length, feature_dim)),通过Reshape将维度调整为(batch_size, 1, 1024),确保Conv1D可以正常处理。

内容的提问来源于stack exchange,提问作者leakie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 15:28:11