Bi-LSTM输出接入Conv1D报错:输入维度不匹配问题求助
问题原因与解决方案
错误根源
Conv1D 层要求输入必须是3维张量,形状为 (batch_size, timesteps, features),但你的第二个Bi-LSTM因为设置了 return_sequences=False,输出是2维张量 (None, 128)(None 是batch维度,128是双向LSTM输出的特征数:64*2),维度不匹配导致报错。
两种可行修复方案
方案1:修改Bi-LSTM的返回模式(推荐)
如果业务逻辑允许保留时间步维度,直接将第二个Bi-LSTM的 return_sequences 改为 True,输出会保留时间步维度变成3维张量,可直接输入Conv1D:
def create_model(hp): inputs = Input(name='inputs',shape=[max_len]) embedding_layer = Embedding(vocab_size, dimention, input_length=max_len)(inputs) bilstm_1 = Bidirectional(LSTM(128, return_sequences=True))(embedding_layer) dropout_1 = Dropout(0.5)(bilstm_1) # 修改return_sequences为True bilstm_2 = Bidirectional(LSTM(64, return_sequences=True))(dropout_1) dropout_2 = Dropout(0.5)(bilstm_2) print(bilstm_2.shape) # 此时形状应为 (None, max_len, 128) conv_1 = Conv1D( filters=hp.Int('conv_1_filter', min_value=32, max_value=128, step=16), kernel_size=hp.Choice('convolution_1', values = [2,6]) )(dropout_2) conv_1 = GlobalMaxPooling1D()(conv_1)
方案2:添加Reshape层适配维度
如果必须使用第二个Bi-LSTM的2维输出,通过Reshape层将其转换为3维张量,注意要保证时间步长度不小于Conv1D的kernel_size:
比如将2维张量(None,128)转为(None,128,1)(时间步为128,特征数为1),这样kernel_size=2或6都满足要求:
def create_model(hp): inputs = Input(name='inputs',shape=[max_len]) embedding_layer = Embedding(vocab_size, dimention, input_length=max_len)(inputs) bilstm_1 = Bidirectional(LSTM(128, return_sequences=True))(embedding_layer) dropout_1 = Dropout(0.5)(bilstm_1) bilstm_2 = Bidirectional(LSTM(64, return_sequences=False))(dropout_1) dropout_2 = Dropout(0.5)(bilstm_2) # 添加Reshape层转换为3维 reshaped = Reshape((128, 1))(dropout_2) print(reshaped.shape) # 此时形状应为 (None, 128, 1) conv_1 = Conv1D( filters=hp.Int('conv_1_filter', min_value=32, max_value=128, step=16), kernel_size=hp.Choice('convolution_1', values = [2,6]) )(reshaped) # 输入改为reshaped conv_1 = GlobalMaxPooling1D()(conv_1)
你之前Reshape失败的可能原因
如果之前试过Reshape但没成功,大概率是维度顺序错误,比如写成Reshape((1,128)),此时时间步为1,而你的kernel_size选了2或6,会触发“时间步长度小于卷积核大小”的错误,所以要保证Reshape后的第一个非batch维度(时间步)≥kernel_size。
内容的提问来源于stack exchange,提问作者Blue Cheese
相关产品推荐
相关产品推荐

