You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark DecisionTree模型特征重要性结果异常问题求助

Spark决策树模型特征重要性异常问题分析

问题现象

运行Spark DecisionTree模型时,其他环节正常,但调用featureImportances查看特征重要性时出现两个异常:

  • 特征得分总和不为1
  • 已知关键分类特征smoking_status的重要性显示为0

特征重要性输出代码:

list(zip(assembler.getInputCols(), decisionTreeModel.featureImportances))

输出结果:

[('age', 0.1717140615500328),
('bmi', 0.001403579349166339),
('hypertension', 0.0),
('heart_disease', 0.0),
('avg_glucose_level', 0.007257486022061398),
('genderVector', 0.0),
('smokingVector', 0.0)]

完整代码

from pyspark.ml.classification import DecisionTreeClassifier
from pyspark.ml.feature import StringIndexer, OneHotEncoder

gender_indexer = StringIndexer(inputCol='gender', outputCol='genderIndexer')
gender_encoder = OneHotEncoder(inputCol='genderIndexer', outputCol='genderVector')

smoking_indexer = StringIndexer(inputCol='smoking_status', outputCol='smokingIndexer')
smoking_encoder = OneHotEncoder(inputCol='smokingIndexer', outputCol='smokingVector')

from pyspark.ml.feature import VectorAssembler
assembler = VectorAssembler(inputCols=['age', 'bmi', 'hypertension', 'heart_disease', 'avg_glucose_level', 'genderVector', 'smokingVector' ], outputCol='features')

classifier = DecisionTreeClassifier(labelCol='stroke', featuresCol='features')

from pyspark.ml import Pipeline
pipeline = Pipeline(stages=[gender_indexer, gender_encoder, smoking_indexer, smoking_encoder, assembler, classifier])

train_data, test_data = strokes.randomSplit([0.7, 0.3])
predictStrokeModel = pipeline.fit(train_data)

result = predictStrokeModel.transform(test_data)

from pyspark.ml.evaluation import MulticlassClassificationEvaluator

evaluator = MulticlassClassificationEvaluator(labelCol='stroke', predictionCol='prediction', metricName='accuracy')
accuracy = evaluator.evaluate(result)

decisionTreeModel = predictStrokeModel.stages[5]
decisionTreeModel.depth

decisionTreeModel.toDebugString

list(zip(assembler.getInputCols(), decisionTreeModel.featureImportances))

异常原因解析

1. 特征得分总和不为1的核心原因

Spark的DecisionTreeModel.featureImportances是按特征维度(而非特征列)计算的:

  • 数值特征(如age、bmi)每个列对应1个维度
  • OneHotEncoder生成的向量列(如genderVector、smokingVector)每个类别对应1个维度(默认丢弃最后一个类别避免共线性)

你用zip(assembler.getInputCols(), ...)把整个向量列和单个重要性值绑定,相当于把向量列下所有维度的重要性总和当成了一个值,这是错误的统计方式。实际featureImportances的长度等于features向量的总维度数,只有被用于分裂的维度重要性非零,且这些非零值的总和为1。

2. smoking_status重要性显示为0的原因

  • 统计方式错误:smokingVector对应多个维度,你现在的统计是把这些维度的重要性全部加总后显示为0,但可能其中部分维度有非零值,只是总和被误算为0。
  • 模型未选择该特征分裂:如果smoking_status编码后的所有维度都没被选为分裂特征,会出现全0的情况,可能的触发因素包括:
    • 训练数据中smoking_status与标签stroke的相关性极低
    • 该特征存在大量缺失值,被模型自动忽略
    • 决策树深度设置过小,模型优先选择了age这类信息增益更高的特征,没有机会用到smoking_status
    • 类别分布极端(如某类别占比99%),导致该特征的区分度不足

修复方案

  1. 正确匹配特征维度与重要性:展开OneHotEncoder的每个维度,生成对应的特征名后再关联重要性:
    # 构建完整的特征维度名称列表
    feature_names = []
    # 添加数值特征列
    numeric_cols = ['age', 'bmi', 'hypertension', 'heart_disease', 'avg_glucose_level']
    feature_names.extend(numeric_cols)
    # 解析genderVector的维度名称
    gender_indexer_model = predictStrokeModel.stages[0]
    for cat in gender_indexer_model.labels[:-1]:
        feature_names.append(f'gender_{cat}')
    # 解析smokingVector的维度名称
    smoking_indexer_model = predictStrokeModel.stages[2]
    for cat in smoking_indexer_model.labels[:-1]:
        feature_names.append(f'smoking_status_{cat}')
    # 关联维度名称与重要性
    importance_list = list(zip(feature_names, decisionTreeModel.featureImportances))
    print(importance_list)
    
  2. 检查树结构:查看decisionTreeModel.toDebugString的输出,确认模型是否真的未使用smoking_status相关维度进行分裂。
  3. 数据与模型调优:
    • 检查smoking_status列的数据质量(缺失值、类别分布)
    • 调整决策树参数:增大maxDepth、减小minInstancesPerNode,给模型更多机会使用次要特征
    • 尝试不做OneHotEncoder,直接用StringIndexer的结果作为序数特征输入(仅适用于有序分类场景)

内容的提问来源于stack exchange,提问作者Renata Mesquita

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 19:23:18