Python OOP实现ROC指标表时列名重复追加'predicted'问题求助
解决Python OOP中重复生成"predicted"前缀列的问题
我一眼就看出问题出在哪了——你直接在传入的原始DataFrame上修改数据,而且没有做任何隔离,再加上方法本身还有几个小bug,导致每次调用都会重复生成带predicted_前缀的列。咱们一步步拆解问题,再给你修复好:
问题根源
- 直接修改原始DataFrame:DataFrame是可变对象,你在
viewer()里直接给data新增列,会永久改变传入的原始数据。下次实例化或者调用方法时,循环data.iloc[:,1:]会把之前生成的predicted_xxx列也包含进去,进而生成predicted_predicted_xxx这种重复前缀的列。 viewer()方法缺少self参数:作为类的实例方法,必须把self作为第一个参数,不然没法访问实例的self.data属性。- 未定义的
threshold变量:代码里用到了threshold但没赋值,运行时会直接报错。 - 列计数时机错误:
count_algo = len(data.columns)是在生成新列之前计算的,但如果原始数据已经有旧的predicted_列,这个计数就会包含这些冗余列,导致后续逻辑混乱。
修复方案
针对这些问题,我们可以做以下调整:
- 操作DataFrame的副本,避免修改原始数据
- 修正实例方法的参数,正确使用
self - 明确传入或定义
threshold变量 - 只处理原始的模型预测列(排除
actual_label和已有的predicted_列)
修正后的完整代码
import pandas as pd from sklearn.metrics import roc_auc_score, accuracy_score, cohen_kappa_score, recall_score, precision_score, f1_score class roc_table: def __init__(self, data, threshold=0.5): # 初始化时创建数据副本,避免修改原始数据 self.data = data.copy() self.threshold = threshold # 把阈值作为实例属性,支持外部传入 def viewer(self): # 只筛选出需要处理的原始预测列:排除actual_label,以及已经带predicted_前缀的列 original_pred_cols = [col for col in self.data.columns if col != 'actual_label' and not col.startswith('predicted_')] # 在副本上生成新的二值化预测列 for col in original_pred_cols: self.data[f'predicted_{col}'] = (self.data[col] >= self.threshold).astype('int') # 提取所有新生成的predicted_列 predicted_cols = [col for col in self.data.columns if col.startswith('predicted_')] # 计算各项指标 metrics_dict = { "AUC": [round(roc_auc_score(self.data['actual_label'], self.data[col]), 2) for col in predicted_cols], "Accuracy": [round(accuracy_score(self.data['actual_label'], self.data[col]), 2) for col in predicted_cols], "Kappa": [round(cohen_kappa_score(self.data['actual_label'], self.data[col]), 2) for col in predicted_cols], "Sensitivity (Recall)": [round(recall_score(self.data['actual_label'], self.data[col]), 2) for col in predicted_cols], "Specificity": [round(1 - recall_score(self.data['actual_label'], self.data[col], pos_label=0), 2) for col in predicted_cols], # 修正:之前的Specificity计算错误,应该是负样本的召回率 "Precision": [round(precision_score(self.data['actual_label'], self.data[col]), 2) for col in predicted_cols], "F1": [round(f1_score(self.data['actual_label'], self.data[col]), 2) for col in predicted_cols] } # 转换为DataFrame并设置列名 rock_table = pd.DataFrame.from_dict(metrics_dict, orient='index').reset_index() col_names = ['metrics'] + [col.replace('predicted_', '') for col in predicted_cols] rock_table.columns = col_names return rock_table
额外说明
- 我还修正了Specificity的计算错误:你之前用了和Accuracy一样的函数,这是不对的,Specificity应该是负样本的召回率,也就是
1 - recall_score(..., pos_label=0)。 - 现在每次实例化或调用
viewer()时,都会基于原始数据的副本操作,不会污染原始数据,也不会重复生成predicted_前缀的列。 - 阈值
threshold现在可以在初始化类的时候传入,默认值设为0.5,使用起来更灵活。
内容的提问来源于stack exchange,提问作者sam c
相关产品推荐
相关产品推荐

