列表索引数组报错:XGBoost网格搜索抽样问题求助
解决XGBoost网格搜索抽样索引错误问题
错误原因
你的代码存在两个核心问题:
- 错误抽取了行数据而非行索引:
random.sample(list(X_res),5000)会将X_res的每一行转为列表元素并抽样,得到的是行内容而非行的位置索引,导致后续索引数组时出错。 - numpy数组不支持直接用Python列表索引,需使用numpy整数数组或布尔数组。
修正方案
方法1:用Python标准库抽样索引并转numpy数组
import random import numpy as np # 抽取5000个不重复的行索引 indices = random.sample(range(len(X_res)), 5000) # 转换为numpy整数数组,适配numpy索引规则 indices_np = np.array(indices) # 传入抽样后的样本到网格搜索函数 model_params = xgboost_search(X_res[indices_np], Y_res[indices_np])
方法2:用numpy随机抽样直接生成索引数组(更高效)
import numpy as np # 直接生成5000个不重复的行索引 indices_np = np.random.choice(len(X_res), size=5000, replace=False) # 传入抽样后的样本 model_params = xgboost_search(X_res[indices_np], Y_res[indices_np])
可选优化:简化网格搜索的交叉验证参数
你的GridSearchCV中cv=skf.split(X,y)可以简化为cv=skf,GridSearchCV会自动处理数据拆分,无需手动传入split结果:
grid_search = GridSearchCV(estimator=xgb, param_grid=params, scoring="roc_auc", n_jobs=1, cv=skf, verbose=search_verbose)
内容的提问来源于stack exchange,提问作者Ladi Latona
相关产品推荐
相关产品推荐

