如何向Embedding层输入列表?多输入拼接维度错误排查
解决Embedding层处理多元素列表输入的ValueError问题
嘿,我帮你分析下你遇到的这个ValueError: total size of new array must be unchanged问题——根源出在你对tags输入的Embedding输出做Reshape的方式上。
先拆解下你的代码逻辑:
- user和movie输入都是
shape=(1,),经过Embedding后得到(batch_size,1,n_factors),再用Reshape((n_factors,))把多余的维度去掉,变成(batch_size,n_factors),这部分完全没问题。 - 但tags输入是
shape=(25,),经过Embedding层后,输出的形状是(batch_size,25,n_factors)——这时候你直接用Reshape((n_factors,)),相当于要把25*n_factors个元素压缩成n_factors个,总元素数直接变了,自然会触发报错。
两种可行的修改方案
方案一:用全局平均池化合并25个Embedding向量
这是最常用的方式,把25个长度为n_factors的向量做平均,得到一个统一的n_factors维度向量,和user、movie的输出维度对齐:
tags = Input(shape=(25,)) t = Embedding(500, n_factors)(tags) # 替换原来的Reshape,用全局平均池化 t = GlobalAveragePooling1D()(t) # 输出形状: (batch_size, n_factors)
方案二:展平后用Dense层压缩维度
如果想让模型学习更复杂的组合方式,可以先把25个Embedding向量展平成一个长向量,再通过全连接层压缩到n_factors维度:
tags = Input(shape=(25,)) t = Embedding(500, n_factors)(tags) t = Flatten()(t) # 输出形状: (batch_size, 25*n_factors) t = Dense(n_factors, activation='relu')(t) # 输出形状: (batch_size, n_factors)
修改后的完整代码示例
from tensorflow.keras.layers import Input, Embedding, Reshape, Concatenate, Dropout, Dense, Activation, Lambda, GlobalAveragePooling1D from tensorflow.keras.regularizers import l2 from tensorflow.keras.optimizers import Adam from tensorflow.keras.models import Model def rec(n_users, n_movies, n_factors, min_rating, max_rating): user = Input(shape=(1,)) u = Embedding(n_users, n_factors, embeddings_initializer='he_normal', embeddings_regularizer=l2(1e-6))(user) u = Reshape((n_factors,))(u) movie = Input(shape=(1,)) m = Embedding(n_movies, n_factors, embeddings_initializer='he_normal', embeddings_regularizer=l2(1e-6))(movie) m = Reshape((n_factors,))(m) tags = Input(shape=(25,)) t = Embedding(500, n_factors)(tags) # 使用方案一的全局平均池化 t = GlobalAveragePooling1D()(t) x = Concatenate()([u, m, t]) x = Dropout(0.05)(x) x = Dense(10, kernel_initializer='he_normal')(x) x = Activation('relu')(x) x = Dropout(0.5)(x) x = Dense(1, kernel_initializer='he_normal')(x) x = Activation('sigmoid')(x) x = Lambda(lambda x: x * (max_rating - min_rating) + min_rating)(x) model = Model(inputs=[user, movie, tags], outputs=x) opt = Adam(lr=0.001) model.compile(loss='mean_squared_error', optimizer=opt) return model
另外确认下你的X_train_array结构是没问题的:第三个元素是形状为(样本数,25)的数组,完全匹配Input(shape=(25,))的要求,修改后应该就能正常运行了。
内容的提问来源于stack exchange,提问作者ashishjohn
相关产品推荐
相关产品推荐

