You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

H2O中如何获取原始类别特征的特征重要性

问题

当数据集包含类别特征时,H2O会自动执行独热编码(one-hot encoding)并启动训练,但调用summary方法查看特征重要性时,会将每个编码后的子特征视为独立特征。如何获取原始类别特征(如示例中的列“0”)的特征重要性?

示例代码

# import libraries
import pandas as pd
import h2o
import random
from h2o.estimators.glm import H2OGeneralizedLinearEstimator

# initiate h20
h2o.init(ip='localhost')   
h2o.remove_all()  

# load a fake data
training_data = h2o.import_file("http://h2o-public-test-data.s3.amazonaws.com/smalldata/glm_test/gamma_dispersion_factor_9_10kRows.csv")

# Specify the predictors (x) and the response (y). Add a dummy categorical column named "0"
myPredictors = ["abs.C1.", "abs.C2.", "abs.C3.", "abs.C4.", "abs.C5.", '0']
myResponse = "resp"

# add a dummy column consisting of random string values
train = h2o.as_list(training_data)
train = pd.concat([train, pd.DataFrame(random.choices(['ya','ne','agh','c','de'], k=len(training_data)))], axis=1)
train = h2o.H2OFrame(train)

# define linear regression method
def linearRegression(df, predictors, response):
    model = H2OGeneralizedLinearEstimator(family="gaussian", lambda_=0, standardize=True)
    model.train(x=predictors, y=response, training_frame=df)
    print(model.summary)

linearRegression(train, myPredictors, myResponse)   

当前输出结果

Variable Importances: 
variable    relative_importance scaled_importance   percentage
0   abs.C5. 1.508031    1.000000    0.257004
1   abs.C4. 1.364653    0.904924    0.232569
2   abs.C3. 1.158184    0.768011    0.197382
3   abs.C2. 0.766653    0.508380    0.130656
4   abs.C1. 0.471997    0.312989    0.080440
5   0.de    0.275667    0.182799    0.046980
6   0.ne    0.210085    0.139311    0.035803
7   0.ya    0.078100    0.051789    0.013310
8   0.c 0.034353    0.022780    0.005855

注:示例仅为最小可复现示例,实际场景中存在更多类别特征。


解决方案

要将编码后的子特征重要性汇总到原始类别特征,可通过以下方式实现:

方法一:手动汇总特征重要性

这是通用方案,适用于所有H2O模型,步骤如下:

  1. 从模型中提取特征重要性数据并转为Pandas DataFrame
  2. 拆分编码后的特征名称,提取原始类别特征的前缀
  3. 按原始特征分组,汇总重要性指标并重新计算归一化值

示例代码:

def get_original_feature_importance(model):
    # 获取特征重要性数据
    var_importance = model.varimp(use_pandas=True)
    # 提取原始特征名称(拆分独热编码的后缀)
    var_importance['original_feature'] = var_importance['variable'].apply(
        lambda x: x.split('.')[0] if '.' in x else x
    )
    # 按原始特征分组汇总
    grouped = var_importance.groupby('original_feature').agg(
        total_relative_importance=('relative_importance', 'sum'),
        total_percentage=('percentage', 'sum')
    ).reset_index()
    # 计算scaled_importance(以最大值为基准归一化)
    max_importance = grouped['total_relative_importance'].max()
    grouped['scaled_importance'] = grouped['total_relative_importance'] / max_importance
    # 按重要性降序排序并调整列顺序
    grouped = grouped.sort_values('total_relative_importance', ascending=False)
    grouped = grouped[['original_feature', 'total_relative_importance', 'scaled_importance', 'total_percentage']]
    return grouped

# 训练模型后调用函数获取原始特征重要性
model = H2OGeneralizedLinearEstimator(family="gaussian", lambda_=0, standardize=True)
model.train(x=myPredictors, y=myResponse, training_frame=train)
original_importance = get_original_feature_importance(model)
print(original_importance)

运行后会得到以原始特征(如0)为单位的汇总结果,其中total_relative_importance是该类别下所有子特征重要性之和,total_percentage是对应占比,scaled_importance按最大值归一化。

方法二:模型原生参数(仅部分模型支持)

部分H2O模型(如GBM、XGBoost)支持通过varimp(True)直接返回原始特征的重要性,但GLM模型默认不支持,因此手动汇总仍是更通用的方案。


内容的提问来源于stack exchange,提问作者sergey_208

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 22:15:26