Keras中1D CNN与LSTM层连接:无需TimeDistributed的端到端训练
问题描述
我构建了如下所示的模型:
def og_build_model_5layer(n_rows,n_cols): input = Input(shape=(n_cols,n_rows),name='INP') print('model_input shape:' , input.shape) c1 = Conv1D(50, 3,name = 'conv_1',padding='same',kernel_initializer="glorot_uniform")(input) b1 = BatchNormalization(name = 'BN_1')(c1) a1 = Activation('relu')(b1) c2 = Conv1D(50,3,name = 'conv_2',padding='same',kernel_initializer="glorot_uniform")(a1) b2 = BatchNormalization(name = 'BN_2')(c2) a2 = Activation('relu')(b2) c3 = Conv1D(50, 3,name = 'conv_3',padding='same',kernel_initializer="glorot_uniform")(a2) b3 = BatchNormalization(name = 'BN_3')(c3) a3 = Activation('relu')(b3) c4 = Conv1D(50, 3,name = 'conv_4',padding='same',kernel_initializer="glorot_uniform")(a3) b4 = BatchNormalization(name = 'BN_4')(c4) a4 = Activation('relu')(b4) c5 = Conv1D(50, 3,name = 'conv_5',padding='same',kernel_initializer="glorot_uniform")(a4) b5 = BatchNormalization(name = 'BN_5')(c5) a5 = Activation('relu')(b5) ######## ADD one LSTM layer HERE ################## fl = Flatten(name='fl')(LSTM_OUTPUT) den = Dense(30,name='dense_1')(fl) drp = Dropout(0.5)(den) output = Dense(1, activation='sigmoid')(drp) opt = Adam(learning_rate=1e-4) model = Model(inputs=input, outputs=output, name='model') extractor = Model(inputs=input,outputs = model.get_layer('fl').output) model.compile(optimizer=opt, loss='binary_crossentropy', metrics=['accuracy']) print(model.summary()) return model,extractor
该模型包含5层Conv1D,每层接收单张图像输入。我希望添加一个LSTM层来处理由200张图像组成的序列,并对整个CNN+LSTM模型进行端到端训练。但我困惑的是,LSTM需要序列输入,而之前的Conv1D层仅处理单张输入,且我不想使用TimeDistributed,请问这种端到端训练是否可行?
解决方案
可行,只需调整输入维度和模型结构,让CNN部分直接处理序列化图像输入,无需TimeDistributed包装。具体步骤如下:
1. 修改输入形状
原输入是单张图像的(n_cols, n_rows),现在要处理200张图像的序列,输入形状需改为(200, n_cols, n_rows)——第一个维度是序列长度(200张图像),后两个维度是单张图像的尺寸。
2. 适配CNN层的序列输入
Keras的Conv1D会默认作用在输入的第二个维度上,当输入为(200, n_cols, n_rows)时,它会在n_cols维度做卷积,同时完整保留第一个序列维度(200),输出形状为(200, n_cols, 50),这个输出可以直接传递给LSTM层,完全符合LSTM的序列输入要求。
3. 添加LSTM层并调整后续结构
在CNN输出a5后添加LSTM层,设置合适的隐藏单元数,若只需整个序列的最终输出,设置return_sequences=False即可。调整后的完整模型代码如下:
def build_cnn_lstm_model(n_rows, n_cols, seq_length=200): # 输入形状:(序列长度, 图像cols维度, 图像rows维度) input = Input(shape=(seq_length, n_cols, n_rows), name='INP') print('model_input shape:', input.shape) # CNN部分:直接处理序列化输入,保留序列维度 c1 = Conv1D(50, 3, name='conv_1', padding='same', kernel_initializer="glorot_uniform")(input) b1 = BatchNormalization(name='BN_1')(c1) a1 = Activation('relu')(b1) c2 = Conv1D(50, 3, name='conv_2', padding='same', kernel_initializer="glorot_uniform")(a1) b2 = BatchNormalization(name='BN_2')(c2) a2 = Activation('relu')(b2) c3 = Conv1D(50, 3, name='conv_3', padding='same', kernel_initializer="glorot_uniform")(a2) b3 = BatchNormalization(name='BN_3')(c3) a3 = Activation('relu')(b3) c4 = Conv1D(50, 3, name='conv_4', padding='same', kernel_initializer="glorot_uniform")(a3) b4 = BatchNormalization(name='BN_4')(c4) a4 = Activation('relu')(b4) c5 = Conv1D(50, 3, name='conv_5', padding='same', kernel_initializer="glorot_uniform")(a4) b5 = BatchNormalization(name='BN_5')(c5) a5 = Activation('relu')(b5) # 添加LSTM层处理200长度的序列 lstm = LSTM(64, name='lstm_1')(a5) fl = Flatten(name='fl')(lstm) den = Dense(30, name='dense_1')(fl) drp = Dropout(0.5)(den) output = Dense(1, activation='sigmoid')(drp) opt = Adam(learning_rate=1e-4) model = Model(inputs=input, outputs=output, name='cnn_lstm_model') extractor = Model(inputs=input, outputs=model.get_layer('fl').output) model.compile(optimizer=opt, loss='binary_crossentropy', metrics=['accuracy']) print(model.summary()) return model, extractor
4. 适配训练数据
训练数据的形状需对应输入形状,即每个样本为(200, n_cols, n_rows),批量输入形状为(batch_size, 200, n_cols, n_rows)。
关键说明
- 这种方式无需TimeDistributed,因为CNN层会自动保留输入的序列维度,输出的序列长度和输入一致,可直接传递给LSTM。
- 端到端训练完全可行,CNN和LSTM的所有参数会在训练过程中同步更新,不需要单独预训练CNN。
内容的提问来源于stack exchange,提问作者Kathan Vyas
相关产品推荐
相关产品推荐

