如何在Pandas DataFrame中基于列值比较选择对应预测标签?
Pandas实现手动投票集成:选择高置信度模型的预测标签
问题描述
给定包含两个模型预测结果的Pandas DataFrame(仅保留模型预测不一致的样本):
true_y m1_labels m1_probs_0 m1_probs_1 m2_labels m2_probs_0 m2_probs_1 0 0 0.628205 0.371795 1 0.491648 0.508352 0 0 0.564113 0.435887 1 0.474973 0.525027 0 1 0.463897 0.536103 0 0.660307 0.339693 0 1 0.454559 0.545441 0 0.512349 0.487651 0 0 0.608345 0.391655 1 0.499531 0.500469 0 0 0.816127 0.183873 1 0.456669 0.543331 0 1 0.442693 0.557307 0 0.573354 0.426646 1 0 0.653497 0.346503 1 0.487212 0.512788 0 1 0.392380 0.607620 0 0.627419 0.372581 0 1 0.375816 0.624184 0 0.631532 0.368468
数据集字段说明:
true_y:样本真实标签m1_labels/m2_labels:模型m1/m2的预测硬标签m1_probs_0/m1_probs_1:模型m1对类别0、1的预测概率m2_probs_0/m2_probs_1:模型m2对类别0、1的预测概率
需求:对每行样本,选择自身预测类别对应概率更高的模型的硬标签,实现手动投票集成。例如第1行,m1预测0的概率(0.6282)高于m2预测1的概率(0.5084),因此选择m1的标签0。
解决方案
步骤1:提取模型预测标签对应的置信度
首先为每行样本提取两个模型各自预测标签对应的概率值:
# 获取m1预测标签对应的概率 df['m1_confidence'] = df.apply(lambda row: row[f'm1_probs_{row["m1_labels"]}'], axis=1) # 获取m2预测标签对应的概率 df['m2_confidence'] = df.apply(lambda row: row[f'm2_probs_{row["m2_labels"]}'], axis=1)
步骤2:根据置信度选择集成标签
比较两个模型的置信度,选择置信度更高的模型的标签;若置信度相等,可自定义规则(如优先选择m1):
# 生成最终集成标签 df['ensemble_label'] = df.apply( lambda row: row['m1_labels'] if row['m1_confidence'] >= row['m2_confidence'] else row['m2_labels'], axis=1 )
完整可运行代码
import pandas as pd # 构造示例数据集 data = { 'true_y': [0,0,0,0,0,0,0,1,0,0], 'm1_labels': [0,0,1,1,0,0,1,0,1,1], 'm1_probs_0': [0.628205,0.564113,0.463897,0.454559,0.608345,0.816127,0.442693,0.653497,0.392380,0.375816], 'm1_probs_1': [0.371795,0.435887,0.536103,0.545441,0.391655,0.183873,0.557307,0.346503,0.607620,0.624184], 'm2_labels': [1,1,0,0,1,1,0,1,0,0], 'm2_probs_0': [0.491648,0.474973,0.660307,0.512349,0.499531,0.456669,0.573354,0.487212,0.627419,0.631532], 'm2_probs_1': [0.508352,0.525027,0.339693,0.487651,0.500469,0.543331,0.426646,0.512788,0.372581,0.368468] } df = pd.DataFrame(data) # 提取置信度 df['m1_confidence'] = df.apply(lambda row: row[f'm1_probs_{row["m1_labels"]}'], axis=1) df['m2_confidence'] = df.apply(lambda row: row[f'm2_probs_{row["m2_labels"]}'], axis=1) # 生成集成标签 df['ensemble_label'] = df.apply( lambda row: row['m1_labels'] if row['m1_confidence'] >= row['m2_confidence'] else row['m2_labels'], axis=1 ) # 查看核心结果 print(df[['true_y', 'm1_labels', 'm2_labels', 'm1_confidence', 'm2_confidence', 'ensemble_label']])
结果验证
运行代码后,ensemble_label列即为最终的集成标签。例如:
- 第1行:
m1_confidence=0.6282 >m2_confidence=0.5084 → 选择m1_labels=0 - 第3行:
m1_confidence=0.5361 <m2_confidence=0.6603 → 选择m2_labels=0
内容的提问来源于stack exchange,提问作者Fredrik
相关产品推荐
相关产品推荐

