You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在TensorFlow的DNNClassifier中对少数类别(Outcome 2)加权?

解决DNNClassifier处理不平衡数据的类别加权问题

我太懂你这种头疼了——当90%-95%的样本都是Outcome 1时,模型根本懒得学习区分两类,直接全预测Outcome 1就能拿到很高的“表面准确率”,但完全漏掉了你真正关心的Outcome 2。好在TensorFlow的Estimator API原生支持类别加权,而且用DNNClassifier就能实现,完全不用碰底层模块,刚好符合你的维护需求。我来给你一步步讲清楚怎么做:

核心思路:给少数类(Outcome 2)加权重

类别加权的本质是让模型在训练时,对少数类的错误惩罚更重。比如Outcome 2只占5%,那我们可以给它设置一个远高于Outcome 1的权重,让模型哪怕为了减少少数类的错误,宁愿多出现一些Outcome 2的假阳性。

具体实现步骤

1. 先计算类别权重

首先统计训练数据中两类的样本数量,然后计算对应权重。常用的计算方式有两种:

  • 方式一:总样本数/(类别数*该类样本数)
  • 方式二:多数类样本数/少数类样本数(更直观,比如Outcome1有950个,Outcome2有50个,权重就是950/50=19)

示例代码:

import pandas as pd
import tensorflow as tf

# 假设你的训练数据存在train_df中,标签列名为'outcome'(0代表Outcome1,1代表Outcome2)
class_counts = train_df['outcome'].value_counts()
total_samples = len(train_df)
num_classes = 2

# 用方式一计算权重
class_weights = {
    0: total_samples / (num_classes * class_counts[0]),
    1: total_samples / (num_classes * class_counts[1])
}

# 或者用方式二(更简单)
# class_weights = {0: 1.0, 1: class_counts[0]/class_counts[1]}

2. 调整输入函数,加入样本权重

需要在输入函数中给每个样本添加对应的权重,并把权重作为一个特征列传入:

def input_fn(data_df, shuffle=True, batch_size=32):
    # 提取所有特征列(排除标签列)
    feature_dict = {col: tf.convert_to_tensor(data_df[col].values) 
                   for col in data_df.columns if col != 'outcome'}
    # 提取标签
    labels = tf.convert_to_tensor(data_df['outcome'].values)
    
    # 给每个样本分配对应的权重
    sample_weights = tf.convert_to_tensor([class_weights[label] for label in data_df['outcome'].values])
    # 把权重加入特征字典
    feature_dict['sample_weight'] = sample_weights
    
    # 构建数据集
    dataset = tf.data.Dataset.from_tensor_slices((feature_dict, labels))
    if shuffle:
        dataset = dataset.shuffle(buffer_size=len(data_df))
    return dataset.batch(batch_size)

3. 创建DNNClassifier时指定权重列

关键是在DNNClassifier的构造参数中加入weight_column,指定我们刚才定义的权重列名:

# 定义特征列(假设你的特征都是数值型,根据实际情况调整)
feature_columns = [tf.feature_column.numeric_column(col) 
                  for col in train_df.columns if col != 'outcome']

# 初始化DNNClassifier
classifier = tf.estimator.DNNClassifier(
    feature_columns=feature_columns,
    hidden_units=[128, 64],  # 根据你的需求调整隐藏层结构
    n_classes=2,
    weight_column='sample_weight'  # 对应输入函数中的权重列名
)

4. 训练与评估

训练时直接用我们修改后的输入函数即可:

# 训练模型
classifier.train(input_fn=lambda: input_fn(train_df), steps=1000)

# 评估时也要用带权重的输入函数,同时重点关注Outcome2的召回率
eval_results = classifier.evaluate(input_fn=lambda: input_fn(test_df, shuffle=False))
print(f"Outcome2召回率: {eval_results['recall_1']}")

重要注意事项

  • 不要只看整体准确率:你的目标是不漏检Outcome2,所以要重点关注Outcome2的召回率(Recall),哪怕整体准确率有所下降也没关系。
  • 调参权重值:一开始可以用我们计算的权重,之后可以尝试调整(比如把Outcome2的权重从20调到30),找到最适合你需求的平衡点——召回率足够高,同时假阳性在可接受范围内。
  • 其他辅助手段:除了加权,你还可以结合过采样少数类(比如SMOTE)或欠采样多数类,但加权是最直接且不需要修改原始数据集的方法,适合你的场景。

我之前在处理医疗检测的不平衡数据时用过这个方法,确实能有效提升少数类的召回率,完全符合你“宁可假阳性也不漏检”的需求。

内容的提问来源于stack exchange,提问作者Abigail Fox

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:05:23