Keras中正则化器与loss_weights冲突问题求助
问题:Keras函数式API中loss_weights与正则化共存导致的形状不匹配错误
我用Keras函数式API构建了一个拼接双网络的模型,同时想使用loss_weights和正则化机制,以下是最小复现代码:
import numpy as np from tensorflow.keras.models import Model from tensorflow.keras.layers import Dense, Dropout, Input, concatenate from tensorflow.keras.regularizers import L2 x1, x2 = np.random.uniform(size=(1000, 4)), np.random.uniform(size=(1000, 4)) y = np.random.uniform(size=(1000, 2)) loss_weights = np.random.randint(2, size=(1000, 2)) in1 = Input(shape=(4,)) d1 = Dense(100)(in1) in2 = Input(shape=(4,)) d2 = Dense(100)(in2) merge = concatenate([d1, d2]) d3 = Dense(2, bias_regularizer=L2(1e-4), kernel_regularizer=L2(1e-4))(d2) model = Model(inputs=[in1, in2], outputs=d3) model.compile(loss='mse', loss_weights=loss_weights) model.fit([x1, x2], y)
运行后出现如下错误:
--------------------------------------------------------------------------- ValueError Traceback (most recent call last) <ipython-input-4-24c2e74ecfe3> in <module> ----> 1 model.fit([x1, x2], y) 1 frames /usr/local/lib/python3.7/dist-packages/keras/engine/training.py in tf__train_function(iterator) 13 try: 14 do_return = True ---> 15 retval_ = ag__.converted_call(ag__.ld(step_function), (ag__.ld(self), ag__.ld(iterator)), None, fscope) 16 except: 17 do_return = False ValueError: in user code: File "/usr/local/lib/python3.7/dist-packages/keras/engine/training.py", line 1051, in train_function * return step_function(self, iterator) File "/usr/local/lib/python3.7/dist-packages/keras/engine/training.py", line 1040, in step_function ** outputs = model.distribute_strategy.run(run_step, args=(data,)) File "/usr/local/lib/python3.7/dist-packages/keras/engine/training.py", line 1030, in run_step ** outputs = model.train_step(data) File "/usr/local/lib/python3.7/dist-packages/keras/engine/training.py", line 890, in train_step loss = self.compute_loss(x, y, y_pred, sample_weight) File "/usr/local/lib/python3.7/dist-packages/keras/engine/training.py", line 949, in compute_loss y, y_pred, sample_weight, regularization_losses=self.losses) File "/usr/local/lib/python3.7/dist-packages/keras/engine/compile_utils.py", line 238, in __call__ total_loss_metric_value = tf.add_n(loss_metric_values) ValueError: Shapes must be equal rank, but are 2 and 0 From merging shape 0 with other shapes. for '{{node AddN}} = AddN[N=2, T=DT_FLOAT](mul_1, dense_2/kernel/Regularizer/mul)' with input shapes: [1000,2], [].
我不想移除正则化(无正则化时代码可正常运行,且正则化对模型性能有帮助),试过一些类似问题的解决方案但无效,请问该如何解决?
错误原因分析
问题出在loss_weights的使用方式上:你传入的是样本级、输出维度级的权重数组(形状(1000,2)),但Keras在计算总损失时,会将加权后的样本损失(形状(1000,2))和标量的正则化损失(形状())相加,两者形状不匹配,导致报错。
Keras的loss_weights参数默认是输出头级的权重(比如多输出模型中给每个输出分配一个权重值),而不是样本级的权重。如果要实现样本级的损失加权,应该用sample_weight参数,而不是loss_weights。
解决方案
方案1:改用sample_weight实现样本级加权
把原来的loss_weights数组作为sample_weight传入model.fit(),同时model.compile()中只保留正则化配置:
import numpy as np from tensorflow.keras.models import Model from tensorflow.keras.layers import Dense, Input, concatenate from tensorflow.keras.regularizers import L2 x1, x2 = np.random.uniform(size=(1000, 4)), np.random.uniform(size=(1000, 4)) y = np.random.uniform(size=(1000, 2)) sample_weights = np.random.randint(2, size=(1000, 2)) in1 = Input(shape=(4,)) d1 = Dense(100)(in1) in2 = Input(shape=(4,)) d2 = Dense(100)(in2) d3 = Dense(2, bias_regularizer=L2(1e-4), kernel_regularizer=L2(1e-4))(d2) model = Model(inputs=[in1, in2], outputs=d3) model.compile(loss='mse') # 将样本权重传入fit model.fit([x1, x2], y, sample_weight=sample_weights)
方案2:自定义损失函数实现加权
如果必须在损失函数层面处理加权,可以自定义一个带权重的MSE损失,同时保留正则化:
import numpy as np import tensorflow as tf from tensorflow.keras.models import Model from tensorflow.keras.layers import Dense, Input, concatenate from tensorflow.keras.regularizers import L2 x1, x2 = np.random.uniform(size=(1000, 4)), np.random.uniform(size=(1000, 4)) y = np.random.uniform(size=(1000, 2)) loss_weights = np.random.randint(2, size=(1000, 2)) # 自定义加权MSE损失 def weighted_mse(y_true, y_pred): weights = tf.convert_to_tensor(loss_weights, dtype=tf.float32) return tf.reduce_mean(weights * tf.square(y_true - y_pred)) in1 = Input(shape=(4,)) d1 = Dense(100)(in1) in2 = Input(shape=(4,)) d2 = Dense(100)(in2) d3 = Dense(2, bias_regularizer=L2(1e-4), kernel_regularizer=L2(1e-4))(d2) model = Model(inputs=[in1, in2], outputs=d3) model.compile(loss=weighted_mse) model.fit([x1, x2], y)
注意事项
loss_weights的正确用法是给多输出模型的每个输出头分配一个标量权重,比如loss_weights=[0.3, 0.7]对应两个输出头的损失权重。- 样本级的损失加权必须用
sample_weight,它支持形状与输出匹配的数组((样本数, 输出维度)),Keras会自动处理加权后的损失与正则化损失的相加。
内容的提问来源于stack exchange,提问作者akra1
相关产品推荐
相关产品推荐

