You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将决策树及随机森林输出的特征编号转换为真实特征名称?

将随机森林特征编号转换为真实特征名称的方法

嘿,我来帮你搞定这个问题!你现在拿到的这个特征编号数组里,每个数字其实是训练数据中特征的索引,而-2代表这个节点是叶子节点,没有用到特征分裂。要把它们转换成真实特征名称,其实就是做个简单的索引映射,步骤很清晰:

第一步:准备你的特征名称列表

首先你得有训练数据对应的真实特征名称。如果你的训练数据是用Pandas DataFrame(比如X_train)训练的,那直接取它的列名就行:

# 假设X_train是你的训练特征DataFrame
feature_names = X_train.columns.tolist()

如果是普通的numpy数组,那你应该自己维护了一个特征名称的列表,比如:

feature_names = ["特征1", "特征2", "特征3", ...]  # 严格按训练数据的特征顺序排列

第二步:转换特征编号为名称

接下来就可以遍历你的特征数组,把每个编号映射成对应的名称。对于-2的叶子节点,我们可以用一个明确的标识(比如"Leaf Node")来代替:

import numpy as np

# 你的特征数组
feature_array = np.array([41, 0, 0, -2, -2, 55, -2, -2, 40, 45, -2, -2, 44, -2, -2], dtype=np.int64)

# 定义映射函数
def map_feature_id_to_name(feature_id):
    if feature_id == -2:
        return "Leaf Node"
    else:
        return feature_names[feature_id]

# 批量转换为名称数组
feature_names_array = np.vectorize(map_feature_id_to_name)(feature_array)

执行完之后,feature_names_array就是你想要的、包含真实特征名称的数组了。

整合到你的现有代码里

你可以把这段转换逻辑直接加到你现有的随机森林代码后面,完整示例大概是这样:

from sklearn.ensemble import RandomForestRegressor
import numpy as np

# 假设X_train是你的训练特征数据,y_train是对应标签
rf = RandomForestRegressor(n_estimators=100, max_depth=3)
rf.fit(X_train, y_train)  # 别忘了先训练模型!

n_nodes = rf.estimators_[0].tree_.node_count
children_left = rf.estimators_[0].tree_.children_left
children_right = rf.estimators_[0].tree_.children_right
feature = rf.estimators_[0].tree_.feature
threshold = rf.estimators_[0].tree_.threshold
node_depth = np.zeros(shape=n_nodes, dtype=np.int64)
is_leaves = np.zeros(shape=n_nodes, dtype=bool)
stack = [(0, -1)] # seed is the root node id and its parent depth
while len(stack) > 0:
    node_id, parent_depth = stack.pop()
    node_depth[node_id] = parent_depth + 1
    # If we have a test node
    if (children_left[node_id] != children_right[node_id]):
        stack.append((children_left[node_id], parent_depth + 1))
        stack.append((children_right[node_id], parent_depth + 1))
    else:
        is_leaves[node_id] = True

# 开始转换特征编号为名称
feature_names = X_train.columns.tolist()
def map_feature_id_to_name(feature_id):
    return "Leaf Node" if feature_id == -2 else feature_names[feature_id]

feature_names_array = np.vectorize(map_feature_id_to_name)(feature)
print(feature_names_array)

小提示

  • 一定要保证feature_names的顺序和你训练数据的特征顺序完全一致,不然映射出来的名称会对应错误!
  • 如果你的随机森林有多个estimator(比如你这里的100个),上面的代码只处理了第一个树的特征,如果你要处理所有树的特征,只需要循环遍历rf.estimators_就行。

内容的提问来源于stack exchange,提问作者leskovecg

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 00:03:11