NLP任务中1D-CNN接入LSTM的位置及报错解决方法
我在NLP任务中尝试将LSTM加入1D-CNN以提升效果,但不确定LSTM的放置位置。参考资料显示CNN-LSTM模型可将CNN层置于前端,后接LSTM层及输出Dense层,但按此结构搭建模型时出现错误:
ValueError: Input 0 of layer "lstm_4" is incompatible with the layer: expected ndim=3, found ndim=2. Full shape received: (None, 128)
该错误因LSTM期望3D输入数组,但当前输入为2D。请问能否在此位置修复该错误,还是需要调整LSTM的放置位置?
模型代码
from keras.models import Sequential from keras.layers import Input, Embedding, Dense, GlobalMaxPooling1D, Conv2D, MaxPool2D, LSTM, Bidirectional, Lambda, Conv1D, MaxPooling1D, GlobalMaxPooling1D model_lstm = Sequential() model_lstm.add( Embedding(vocab_size ,embed_size ,weights = [embedding_matrix] #Supplied embedding matrix created from glove ,input_length = maxlen ,trainable=False) ) model_lstm.add(SpatialDropout1D(rate = 0.4)) model_lstm.add(Conv1D(256, 7, activation="relu")) model_lstm.add(MaxPooling1D()) #model_lstm.add(LSTM(128, dropout=0.3, recurrent_dropout=0.3, return_sequences=True)) model_lstm.add(Conv1D(128, 5, activation="relu")) model_lstm.add(MaxPooling1D()) model_lstm.add(GlobalMaxPooling1D()) model_lstm.add(LSTM(128, dropout=0.3,return_sequences=True)) model_lstm.add(Dropout(0.3)) model_lstm.add(Dense(128, activation="relu")) model_lstm.add(Dense(4, activation='softmax')) print(model_lstm.summary())
完整预处理代码
print("Train shape : ",train_X2.shape) print("Test shape : ",test_X2.shape) ## Tokenize the sentences tokenizer = Tokenizer(num_words=num_unique_words) tokenizer.fit_on_texts(list(train_X2)) train_X2 = tokenizer.texts_to_sequences(train_X2) test_X2 = tokenizer.texts_to_sequences(test_X2) ## Pad the sentences train_X = pad_sequences(train_X2, maxlen=maxlen) test_X = pad_sequences(test_X2, maxlen=maxlen) word_index = tokenizer.word_index vocab_size = len(tokenizer.word_index) + 1 from sklearn.preprocessing import LabelEncoder from tensorflow.keras.utils import to_categorical #label encoding le = LabelEncoder() train_y = le.fit_transform(train_y2.tolist()) test_y = le.transform(test_y2.tolist()) #one hot encoding train_y = to_categorical(train_y) test_y = to_categorical(test_y) # Word2Vec as pretrained embedding import gensim from gensim.models import Word2Vec from gensim.utils import simple_preprocess from gensim.models.keyedvectors import KeyedVectors NUM_WORDS=20000 word_vectors = KeyedVectors.load_word2vec_format(r'./input/GoogleNews-vectors-negative300.bin', binary=True) EMBEDDING_DIM=300 vocabulary_size=min(len(word_index)+1,NUM_WORDS) embedding_matrix = np.zeros((vocabulary_size, EMBEDDING_DIM)) for word, i in word_index.items(): if i>=NUM_WORDS: continue try: embedding_vector = word_vectors[word] embedding_matrix[i] = embedding_vector except KeyError: embedding_matrix[i]=np.random.normal(0,np.sqrt(0.25),EMBEDDING_DIM) del(word_vectors) from keras.layers import Embedding embedding_layer = Embedding(vocabulary_size, EMBEDDING_DIM, weights=[embedding_matrix], trainable=True) from keras.layers import Embedding EMBEDDING_DIM=300 vocabulary_size=min(len(word_index)+1,NUM_WORDS) embedding_layer = Embedding(vocabulary_size, EMBEDDING_DIM) # CNN
错误原因与修复方案
错误根源
代码中的GlobalMaxPooling1D()层会将CNN输出的3D张量(形状为(batch_size, sequence_length, filters))压缩为2D张量((batch_size, filters)),而LSTM要求输入必须是3D格式((batch_size, timesteps, features)),因此直接在GlobalMaxPooling1D()后添加LSTM会触发维度不兼容错误。
两种修复思路
思路1:调整LSTM位置(推荐,符合标准CNN-LSTM结构)
将LSTM放在CNN之后、GlobalMaxPooling1D()之前,这样CNN输出的3D序列张量可以直接输入LSTM,保留序列信息供LSTM捕捉时序依赖。若LSTM后需接其他序列层,设置return_sequences=True;若直接接Dense层,设为False即可。
修改后的模型代码示例:
model_lstm = Sequential() model_lstm.add( Embedding(vocab_size ,embed_size ,weights = [embedding_matrix] ,input_length = maxlen ,trainable=False) ) model_lstm.add(SpatialDropout1D(rate = 0.4)) model_lstm.add(Conv1D(256, 7, activation="relu")) model_lstm.add(MaxPooling1D()) model_lstm.add(Conv1D(128, 5, activation="relu")) model_lstm.add(MaxPooling1D()) # 将LSTM移至GlobalMaxPooling1D之前 model_lstm.add(LSTM(128, dropout=0.3, return_sequences=False)) model_lstm.add(Dropout(0.3)) model_lstm.add(Dense(128, activation="relu")) model_lstm.add(Dense(4, activation='softmax')) print(model_lstm.summary())
思路2:在GlobalMaxPooling1D后恢复3D维度
若必须保留当前层顺序,可通过Reshape层将2D张量重新转为3D,人为新增时间步维度供LSTM接收,但此方式无实际序列意义,仅为兼容输入格式:
from keras.layers import Reshape # ... 原代码至GlobalMaxPooling1D之后 model_lstm.add(GlobalMaxPooling1D()) # 新增时间步维度,将形状从(None,128)转为(None,1,128) model_lstm.add(Reshape((1, 128))) model_lstm.add(LSTM(128, dropout=0.3, return_sequences=True)) # LSTM输出后需再次压缩维度才能接Dense层 model_lstm.add(GlobalMaxPooling1D()) model_lstm.add(Dropout(0.3)) model_lstm.add(Dense(128, activation="relu")) model_lstm.add(Dense(4, activation='softmax'))
补充说明
标准CNN-LSTM的设计逻辑是:用CNN提取文本局部特征,再将这些特征序列输入LSTM捕捉全局时序关联,因此LSTM必须放在能保留序列维度的CNN输出之后、全局池化层之前,才能发挥作用。
内容的提问来源于stack exchange,提问作者Test

