H2O库Python环境下网格搜索时ignored_columns参数失效问题
H2O网格搜索中
ignored_columns参数失效问题 ignored_columns参数用于在构建模型时指定需要忽略的特征列。
单模型训练时参数正常生效
训练单个随机森林模型时,指定忽略列c,从特征重要性结果可以看到,列c确实未被用于训练:
import pandas as pd import h2o from h2o.estimators import H2ODeepLearningEstimator from h2o.grid.grid_search import H2OGridSearch from h2o.estimators.random_forest import H2ORandomForestEstimator h2o.init() x = pd.DataFrame([[0, 1, 4], [5, 1, 6], [15, 2, 0], [25, 5 , 32], [35, 11 ,89], [45, 15, 1], [55, 34,3], [60, 35,4]], columns = ['a','b','c']) y = pd.DataFrame([4, 5, 20, 14, 32, 22, 38, 43], columns = ['label']) hf = h2o.H2OFrame( pd.concat([x,y], axis="columns")) X = hf.col_names[:-1] y = hf.col_names[-1] model= H2ORandomForestEstimator(ignored_columns = ['c']) model.train(y = y, training_frame=hf) model.varimp(use_pandas=True)
输出结果:
| variable | relative_importance | scaled_importance | percentage |
|---|---|---|---|
| b | 33876.328125 | 1.000000 | 0.540893 |
| a | 28753.998047 | 0.848793 | 0.459107 |
网格搜索时参数失效
但开启网格搜索进行超参数调优时,ignored_columns参数并未生效,列c出现在了特征重要性结果中:
params = {'max_depth': list(range(7, 16)), 'sample_rate': [0.8], } criteria = {'strategy': 'RandomDiscrete', 'max_models': 4} grid = H2OGridSearch(model= H2ORandomForestEstimator(ignored_columns = ['c']), search_criteria=criteria, hyper_params=params ) grid.train( y = y, training_frame=hf) best_model = grid.get_grid(sort_by='rmse', decreasing=False)[0] best_model.varimp(use_pandas=True)
输出结果:
| variable | relative_importance | scaled_importance | percentage |
|---|---|---|---|
| a | 33525.109375 | 1.000000 | 0.516545 |
| b | 23314.916016 | 0.695446 | 0.359230 |
| c | 8062.515137 | 0.240492 | 0.124225 |
内容的提问来源于stack exchange,提问作者whitepanda
相关产品推荐
相关产品推荐

