You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow自定义损失函数如何传入额外样本权重数据?

解决TensorFlow二分类自定义损失函数传递样本权重的问题

针对你遇到的问题,这里提供两种可行的解决方案,无需将权重混入标签数据中:

方法一:将权重作为模型的额外输入

这种方法把样本权重作为模型的第二个输入,在构建模型时明确接收该输入,然后在损失函数中使用它计算带权重的损失,适合规范的大数据集训练流程。

代码实现

import pandas as pd
import tensorflow as tf
from tensorflow import keras

def custom_loss(y_true, y_pred, weights):
    # 将预测概率转为二分类结果(0或1)
    y_pred_class = tf.cast(y_pred > 0.5, tf.float32)
    # 判断预测是否正确
    correct_pred = tf.equal(y_true, y_pred_class)
    
    # 按照需求计算损失:正确预测时用权重*-1,错误时为1
    loss = tf.where(correct_pred, weights * -1.0, tf.ones_like(y_true))
    # 返回平均损失(也可根据需求返回总和)
    return tf.reduce_mean(loss)

exampledata = pd.DataFrame([{"a":1.0,"b":0,"c":0.1,"result":1,"cweight":0.1},
                            {"a":0.8,"b":1,"c":0.5,"result":0,"cweight":0.5},
                            {"a":0.5,"b":1,"c":0.6,"result":0,"cweight":0.5},
                            {"a":0.2,"b":1,"c":0.9,"result":1,"cweight":0.3}])

# 拆分特征、标签和权重,确保维度匹配模型输出
x = exampledata[["a","b","c"]]
y = exampledata["result"].values.reshape(-1, 1)
weights = exampledata["cweight"].values.reshape(-1, 1)

# 构建多输入模型
input_features = keras.layers.Input(shape=(3,), name='features')
input_weights = keras.layers.Input(shape=(1,), name='weights')

x_layer = keras.layers.Dense(256, activation="relu")(input_features)
x_layer = keras.layers.Dropout(0.5)(x_layer)
x_layer = keras.layers.Dense(256, activation="relu")(x_layer)
x_layer = keras.layers.Dropout(0.5)(x_layer)
output = keras.layers.Dense(1, activation="sigmoid")(x_layer)

model = keras.Model(inputs=[input_features, input_weights], outputs=output)

# 包装损失函数,适配Keras默认接口
def loss_wrapper(y_true, y_pred):
    return custom_loss(y_true, y_pred, input_weights)

model.compile(optimizer="adam", loss=loss_wrapper, metrics=["accuracy"])

# 训练时传入特征和权重两个输入
model.fit([x, weights], y, epochs=5)

方法二:用Lambda函数包装损失并传入权重

这种方法不需要修改模型结构,通过嵌套函数将权重捕获到损失函数中,实现起来更简洁。

代码实现

import pandas as pd
import tensorflow as tf
from tensorflow import keras

def custom_loss(weights):
    def loss(y_true, y_pred):
        y_pred_class = tf.cast(y_pred > 0.5, tf.float32)
        correct_pred = tf.equal(y_true, y_pred_class)
        loss = tf.where(correct_pred, weights * -1.0, tf.ones_like(y_true))
        return tf.reduce_mean(loss)
    return loss

exampledata = pd.DataFrame([{"a":1.0,"b":0,"c":0.1,"result":1,"cweight":0.1},
                            {"a":0.8,"b":1,"c":0.5,"result":0,"cweight":0.5},
                            {"a":0.5,"b":1,"c":0.6,"result":0,"cweight":0.5},
                            {"a":0.2,"b":1,"c":0.9,"result":1,"cweight":0.3}])

x = exampledata[["a","b","c"]]
y = exampledata["result"].values.reshape(-1, 1)
# 转换为TensorFlow张量
weights = tf.convert_to_tensor(exampledata["cweight"].values.reshape(-1, 1), dtype=tf.float32)

model = keras.Sequential([
    keras.layers.Dense(256 , input_shape=(3,), activation="relu"),
    keras.layers.Dropout(0.5),
    keras.layers.Dense(256 , activation="relu"),
    keras.layers.Dropout(0.5),
    keras.layers.Dense(1, activation="sigmoid")
])

# 编译时传入带权重的损失函数
model.compile(optimizer="adam", loss=custom_loss(weights), metrics=["accuracy"])

model.fit(x, y, epochs=5)

关键注意事项

  1. 损失逻辑调整:你定义的“正确预测时损失为cweight*-1”,由于TensorFlow默认最小化损失,这会让模型优先降低正确样本的损失。如果需要调整惩罚逻辑,可修改loss计算式,比如正确时用0或正数权重,错误时用更大的惩罚值。
  2. 维度匹配:确保标签y和权重weights的维度与模型输出一致(比如都转为二维数组),避免形状不兼容报错。
  3. 数据集拆分:如果用方法二,训练集和验证集的权重需要分别传入对应的损失函数,不能共用同一组权重。

内容的提问来源于stack exchange,提问作者cmj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 00:55:30