You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Keras中正则化器与loss_weights冲突问题求助

问题:Keras函数式API中loss_weights与正则化共存导致的形状不匹配错误

我用Keras函数式API构建了一个拼接双网络的模型,同时想使用loss_weights和正则化机制,以下是最小复现代码:

import numpy as np
from tensorflow.keras.models import Model
from tensorflow.keras.layers import Dense, Dropout, Input, concatenate
from tensorflow.keras.regularizers import L2

x1, x2 = np.random.uniform(size=(1000, 4)), np.random.uniform(size=(1000, 4))
y = np.random.uniform(size=(1000, 2))
loss_weights = np.random.randint(2, size=(1000, 2))

in1 = Input(shape=(4,))
d1 = Dense(100)(in1)
in2 = Input(shape=(4,))
d2 = Dense(100)(in2)
merge = concatenate([d1, d2])
d3 = Dense(2, bias_regularizer=L2(1e-4), kernel_regularizer=L2(1e-4))(d2)

model = Model(inputs=[in1, in2], outputs=d3)
model.compile(loss='mse', loss_weights=loss_weights)
model.fit([x1, x2], y)

运行后出现如下错误:

---------------------------------------------------------------------------

ValueError                                Traceback (most recent call last)

<ipython-input-4-24c2e74ecfe3> in <module>
----> 1 model.fit([x1, x2], y)

1 frames

/usr/local/lib/python3.7/dist-packages/keras/engine/training.py in tf__train_function(iterator)
     13                 try:
     14                     do_return = True
---> 15                     retval_ = ag__.converted_call(ag__.ld(step_function), (ag__.ld(self), ag__.ld(iterator)), None, fscope)
     16                 except:
     17                     do_return = False

ValueError: in user code:

    File "/usr/local/lib/python3.7/dist-packages/keras/engine/training.py", line 1051, in train_function  *
        return step_function(self, iterator)
    File "/usr/local/lib/python3.7/dist-packages/keras/engine/training.py", line 1040, in step_function  **
        outputs = model.distribute_strategy.run(run_step, args=(data,))
    File "/usr/local/lib/python3.7/dist-packages/keras/engine/training.py", line 1030, in run_step  **
        outputs = model.train_step(data)
    File "/usr/local/lib/python3.7/dist-packages/keras/engine/training.py", line 890, in train_step
        loss = self.compute_loss(x, y, y_pred, sample_weight)
    File "/usr/local/lib/python3.7/dist-packages/keras/engine/training.py", line 949, in compute_loss
        y, y_pred, sample_weight, regularization_losses=self.losses)
    File "/usr/local/lib/python3.7/dist-packages/keras/engine/compile_utils.py", line 238, in __call__
        total_loss_metric_value = tf.add_n(loss_metric_values)

    ValueError: Shapes must be equal rank, but are 2 and 0
        From merging shape 0 with other shapes. for '{{node AddN}} = AddN[N=2, T=DT_FLOAT](mul_1, dense_2/kernel/Regularizer/mul)' with input shapes: [1000,2], [].

我不想移除正则化(无正则化时代码可正常运行,且正则化对模型性能有帮助),试过一些类似问题的解决方案但无效,请问该如何解决?


错误原因分析

问题出在loss_weights的使用方式上:你传入的是样本级、输出维度级的权重数组(形状(1000,2)),但Keras在计算总损失时,会将加权后的样本损失(形状(1000,2))和标量的正则化损失(形状())相加,两者形状不匹配,导致报错。

Keras的loss_weights参数默认是输出头级的权重(比如多输出模型中给每个输出分配一个权重值),而不是样本级的权重。如果要实现样本级的损失加权,应该用sample_weight参数,而不是loss_weights。


解决方案

方案1:改用sample_weight实现样本级加权

把原来的loss_weights数组作为sample_weight传入model.fit(),同时model.compile()中只保留正则化配置:

import numpy as np
from tensorflow.keras.models import Model
from tensorflow.keras.layers import Dense, Input, concatenate
from tensorflow.keras.regularizers import L2

x1, x2 = np.random.uniform(size=(1000, 4)), np.random.uniform(size=(1000, 4))
y = np.random.uniform(size=(1000, 2))
sample_weights = np.random.randint(2, size=(1000, 2))

in1 = Input(shape=(4,))
d1 = Dense(100)(in1)
in2 = Input(shape=(4,))
d2 = Dense(100)(in2)
d3 = Dense(2, bias_regularizer=L2(1e-4), kernel_regularizer=L2(1e-4))(d2)

model = Model(inputs=[in1, in2], outputs=d3)
model.compile(loss='mse')
# 将样本权重传入fit
model.fit([x1, x2], y, sample_weight=sample_weights)

方案2:自定义损失函数实现加权

如果必须在损失函数层面处理加权,可以自定义一个带权重的MSE损失,同时保留正则化:

import numpy as np
import tensorflow as tf
from tensorflow.keras.models import Model
from tensorflow.keras.layers import Dense, Input, concatenate
from tensorflow.keras.regularizers import L2

x1, x2 = np.random.uniform(size=(1000, 4)), np.random.uniform(size=(1000, 4))
y = np.random.uniform(size=(1000, 2))
loss_weights = np.random.randint(2, size=(1000, 2))

# 自定义加权MSE损失
def weighted_mse(y_true, y_pred):
    weights = tf.convert_to_tensor(loss_weights, dtype=tf.float32)
    return tf.reduce_mean(weights * tf.square(y_true - y_pred))

in1 = Input(shape=(4,))
d1 = Dense(100)(in1)
in2 = Input(shape=(4,))
d2 = Dense(100)(in2)
d3 = Dense(2, bias_regularizer=L2(1e-4), kernel_regularizer=L2(1e-4))(d2)

model = Model(inputs=[in1, in2], outputs=d3)
model.compile(loss=weighted_mse)
model.fit([x1, x2], y)

注意事项

  • loss_weights的正确用法是给多输出模型的每个输出头分配一个标量权重,比如loss_weights=[0.3, 0.7]对应两个输出头的损失权重。
  • 样本级的损失加权必须用sample_weight,它支持形状与输出匹配的数组((样本数, 输出维度)),Keras会自动处理加权后的损失与正则化损失的相加。

内容的提问来源于stack exchange,提问作者akra1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 03:45:35