合并XGBoost与Surprise模型时出现NoneType无drop属性报错如何解决
报错解决方案
错误原因
pandas的DataFrame.drop()方法设置inplace=True参数时,会直接修改原DataFrame对象,返回值为None。你将该返回值赋值给new_df变量后,后续调用new_df.drop()本质是对None对象调用方法,自然触发AttributeError: 'NoneType' object has no attribute 'drop'报错。
另外还有一个隐藏问题:model_1如果直接修改传入的原始df,会导致后续model_2需要用到的studentId、testId列被提前删除,也会引发报错。
修正后的model_1代码
def model_1(df, test_indices=None): # 先做数据副本,避免修改原始传入的df df_copy = df.copy() cols_to_drop = ['testId', 'studentId'] # 去掉inplace=True,直接接收drop后的新DataFrame new_df = df_copy.drop(cols_to_drop, axis=1) X = new_df.drop('result', axis=1) y = new_df['result'] # 如果传入了指定的测试集索引,就用指定的,保证和model2测试集对齐 if test_indices is not None: X_train, X_test = X.drop(test_indices), X.loc[test_indices] y_train, y_test = y.drop(test_indices), y.loc[test_indices] else: X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=5) model = XGBClassifier() model.fit(X_train, y_train) # 输出正类预测概率,和协同过滤的0-1分数值域对齐 y_pred_proba = model.predict_proba(X_test)[:, 1] # 把测试集的主键和预测结果一起返回,方便后续对齐 res_df = df_copy.loc[X_test.index, ['studentId', 'testId', 'result']].copy() res_df['xgb_predictions'] = y_pred_proba return res_df
双模型正确融合方案
融合前需要解决两个核心问题:
- 两个模型必须使用同一份训练集、同一份测试集,避免预测的样本不匹配
- 两个模型的预测结果值域要统一,都转为0-1之间的连续概率/分数,再加权融合,最后设置阈值(比如0.5)得到最终分类标签
融合实现代码
def model_2(df, test_indices=None): df_copy = df.copy() reader = Reader(rating_scale=(0, 1)) # 同样按指定索引切分数据集,保证和xgb的测试集一致 if test_indices is not None: train_df = df_copy.drop(test_indices) test_df = df_copy.loc[test_indices] data = Dataset.load_from_df(train_df[['studentId', 'testId', 'result']], reader) trainset = data.build_full_trainset() # 构造surprise格式的测试集 testset = list(test_df[['studentId', 'testId', 'result']].itertuples(index=False, name=None)) else: data = Dataset.load_from_df(df_copy[['studentId', 'testId', 'result']], reader) trainset, testset = train_test_split(data, test_size=0.25, random_state=5) algo = KNNWithMeans() algo.fit(trainset) test = algo.test(testset) test = pd.DataFrame(test) test.drop("details", inplace=True, axis=1) test.columns = ['studentId', 'testId', 'actual', 'cf_predictions'] return test def merged_models(df, xgb_weight=0.5, cf_weight=0.5, threshold=0.5): # 先全局切分训练集和测试集,保证两个模型用同一份数据 total_indices = df.index train_indices, test_indices = train_test_split(total_indices, test_size=0.25, random_state=5) # 分别获取两个模型的预测结果 xgb_res = model_1(df, test_indices) cf_res = model_2(df, test_indices) # 按主键对齐两个预测结果 merge_df = pd.merge(xgb_res, cf_res, on=['studentId', 'testId'], how='inner') # 加权融合 merge_df['final_score'] = xgb_weight * merge_df['xgb_predictions'] + cf_weight * merge_df['cf_predictions'] # 按阈值得到最终分类标签 merge_df['final_pred'] = (merge_df['final_score'] >= threshold).astype(int) return merge_df
混合模型效果评估
得到融合后的预测结果表merge_df后,用真实标签result(和actual是同一个值,任选其一即可)和final_pred/final_score计算分类指标即可:
- 分类硬指标:导入
sklearn.metrics下的accuracy_score、precision_score、recall_score、f1_score,分别传入真实标签和final_pred计算 - 排序/概率指标:导入
sklearn.metrics下的roc_auc_score,传入真实标签和final_score计算
如果需要调整最优权重,可以把权重作为超参数,在验证集上做网格搜索,选择指标最优的权重组合即可。
内容的提问来源于stack exchange,提问作者futuredataengineer
相关产品推荐
相关产品推荐

