You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

NLP任务中1D-CNN接入LSTM的位置及报错解决方法

问题:CNN-LSTM模型输入维度不兼容错误修复

我在NLP任务中尝试将LSTM加入1D-CNN以提升效果,但不确定LSTM的放置位置。参考资料显示CNN-LSTM模型可将CNN层置于前端,后接LSTM层及输出Dense层,但按此结构搭建模型时出现错误:

ValueError: Input 0 of layer "lstm_4" is incompatible with the layer: expected ndim=3, found ndim=2. Full shape received: (None, 128)

该错误因LSTM期望3D输入数组,但当前输入为2D。请问能否在此位置修复该错误,还是需要调整LSTM的放置位置?


模型代码

from keras.models import Sequential
from keras.layers import Input, Embedding, Dense, GlobalMaxPooling1D, Conv2D, MaxPool2D, LSTM, Bidirectional, Lambda, Conv1D, MaxPooling1D, GlobalMaxPooling1D

model_lstm = Sequential()

model_lstm.add(
        Embedding(vocab_size
                ,embed_size
                ,weights = [embedding_matrix] #Supplied embedding matrix created from glove
                ,input_length = maxlen
                ,trainable=False)
         )
model_lstm.add(SpatialDropout1D(rate = 0.4))
model_lstm.add(Conv1D(256, 7, activation="relu"))
model_lstm.add(MaxPooling1D())
#model_lstm.add(LSTM(128, dropout=0.3, recurrent_dropout=0.3, return_sequences=True))
model_lstm.add(Conv1D(128, 5, activation="relu"))
model_lstm.add(MaxPooling1D())
model_lstm.add(GlobalMaxPooling1D())
model_lstm.add(LSTM(128, dropout=0.3,return_sequences=True))
model_lstm.add(Dropout(0.3))
model_lstm.add(Dense(128, activation="relu"))
model_lstm.add(Dense(4, activation='softmax'))
print(model_lstm.summary())

完整预处理代码

print("Train shape : ",train_X2.shape)
print("Test shape : ",test_X2.shape)

## Tokenize the sentences
tokenizer = Tokenizer(num_words=num_unique_words)
tokenizer.fit_on_texts(list(train_X2))
train_X2 = tokenizer.texts_to_sequences(train_X2)
test_X2 = tokenizer.texts_to_sequences(test_X2)

## Pad the sentences 
train_X = pad_sequences(train_X2, maxlen=maxlen)
test_X = pad_sequences(test_X2, maxlen=maxlen)

word_index = tokenizer.word_index
vocab_size = len(tokenizer.word_index) + 1

from sklearn.preprocessing import LabelEncoder
from tensorflow.keras.utils import to_categorical

#label encoding
le = LabelEncoder()
train_y = le.fit_transform(train_y2.tolist())
test_y = le.transform(test_y2.tolist())

#one hot encoding
train_y = to_categorical(train_y)
test_y = to_categorical(test_y)

# Word2Vec as pretrained embedding
import gensim
from gensim.models import Word2Vec
from gensim.utils import simple_preprocess

from gensim.models.keyedvectors import KeyedVectors
NUM_WORDS=20000
word_vectors = KeyedVectors.load_word2vec_format(r'./input/GoogleNews-vectors-negative300.bin', binary=True)

EMBEDDING_DIM=300
vocabulary_size=min(len(word_index)+1,NUM_WORDS)
embedding_matrix = np.zeros((vocabulary_size, EMBEDDING_DIM))
for word, i in word_index.items():
    if i>=NUM_WORDS:
        continue
    try:
        embedding_vector = word_vectors[word]
        embedding_matrix[i] = embedding_vector
    except KeyError:
        embedding_matrix[i]=np.random.normal(0,np.sqrt(0.25),EMBEDDING_DIM)

del(word_vectors)

from keras.layers import Embedding
embedding_layer = Embedding(vocabulary_size,
                            EMBEDDING_DIM,
                            weights=[embedding_matrix],
                            trainable=True)

from keras.layers import Embedding
EMBEDDING_DIM=300
vocabulary_size=min(len(word_index)+1,NUM_WORDS)

embedding_layer = Embedding(vocabulary_size,
                            EMBEDDING_DIM)

# CNN

错误原因与修复方案

错误根源

代码中的GlobalMaxPooling1D()层会将CNN输出的3D张量(形状为(batch_size, sequence_length, filters))压缩为2D张量((batch_size, filters)),而LSTM要求输入必须是3D格式((batch_size, timesteps, features)),因此直接在GlobalMaxPooling1D()后添加LSTM会触发维度不兼容错误。

两种修复思路

思路1:调整LSTM位置(推荐,符合标准CNN-LSTM结构)

将LSTM放在CNN之后、GlobalMaxPooling1D()之前,这样CNN输出的3D序列张量可以直接输入LSTM,保留序列信息供LSTM捕捉时序依赖。若LSTM后需接其他序列层,设置return_sequences=True;若直接接Dense层,设为False即可。

修改后的模型代码示例:

model_lstm = Sequential()

model_lstm.add(
        Embedding(vocab_size
                ,embed_size
                ,weights = [embedding_matrix]
                ,input_length = maxlen
                ,trainable=False)
         )
model_lstm.add(SpatialDropout1D(rate = 0.4))
model_lstm.add(Conv1D(256, 7, activation="relu"))
model_lstm.add(MaxPooling1D())
model_lstm.add(Conv1D(128, 5, activation="relu"))
model_lstm.add(MaxPooling1D())
# 将LSTM移至GlobalMaxPooling1D之前
model_lstm.add(LSTM(128, dropout=0.3, return_sequences=False))
model_lstm.add(Dropout(0.3))
model_lstm.add(Dense(128, activation="relu"))
model_lstm.add(Dense(4, activation='softmax'))
print(model_lstm.summary())

思路2:在GlobalMaxPooling1D后恢复3D维度

若必须保留当前层顺序,可通过Reshape层将2D张量重新转为3D,人为新增时间步维度供LSTM接收,但此方式无实际序列意义,仅为兼容输入格式:

from keras.layers import Reshape

# ... 原代码至GlobalMaxPooling1D之后
model_lstm.add(GlobalMaxPooling1D())
# 新增时间步维度,将形状从(None,128)转为(None,1,128)
model_lstm.add(Reshape((1, 128)))
model_lstm.add(LSTM(128, dropout=0.3, return_sequences=True))
# LSTM输出后需再次压缩维度才能接Dense层
model_lstm.add(GlobalMaxPooling1D())
model_lstm.add(Dropout(0.3))
model_lstm.add(Dense(128, activation="relu"))
model_lstm.add(Dense(4, activation='softmax'))

补充说明

标准CNN-LSTM的设计逻辑是:用CNN提取文本局部特征,再将这些特征序列输入LSTM捕捉全局时序关联,因此LSTM必须放在能保留序列维度的CNN输出之后、全局池化层之前,才能发挥作用。


内容的提问来源于stack exchange,提问作者Test

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 18:50:26