You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Numpy特征组的回归逐步子集选择器构建疑问

回归逐步子集选择器:特征组层面的筛选实现

我明白你现在的需求——之前能搞定单特征列的逐步选择,但现在要升级到按特征组(比如a、b、c这种整组)来做逐步筛选,对吧?其实核心思路就是把操作单元从「单个特征列」换成「整个特征组」,不用直接去碰拼接后的大矩阵,而是先跟踪哪些组在模型里、哪些被排除,每次迭代基于选中的组拼接特征再评估。下面给你具体的实现方案:

1. 先给特征组做标识

首先给每个特征组起个好区分的名字或者用索引,方便我们跟踪状态。比如用名称的话:

import numpy as np
# 你的原始特征组
a = np.array([[ 1., 1.], [ 1., 1.], [ 1., 1.]])
b = np.array([[ 88., 42.5, 9. ], [ 121.5, 76., 42.5], [ 167., 121.5, 88. ]])
c = np.array([[ 88., 42.5, 13. ], [ 117.5, 72., 42.5], [ 163., 117.5, 88. ]])

# 定义特征组的名称和对应的组列表
group_names = ['a', 'b', 'c']
total_features = [a, b, c]

2. 初始化特征组的筛选状态

如果是向前逐步选择(从无到有,每次加最优组),初始状态是模型里没有任何组,所有组都在排除列表里:

features_in_model = []  # 当前模型包含的特征组
excluded = group_names.copy()  # 初始时所有组都待选

如果是向后逐步剔除(从全组开始,每次删性能影响最小的组),初始状态是模型包含所有组,排除列表为空:

features_in_model = group_names.copy()  # 初始包含所有组
excluded = []  # 初始无被排除的组

3. 完整的向前逐步选择实现

下面是带模型评估的完整代码,用R²作为性能指标(你可以换成MSE、MAE等):

from sklearn.linear_model import LinearRegression
from sklearn.metrics import r2_score

# 模拟目标变量(你替换成自己的真实y即可)
y = np.array([10., 20., 30.])

best_score = -np.inf

# 迭代筛选,直到所有组都加入模型
while excluded:
    current_best_score = -np.inf
    best_group = None
    
    # 遍历所有待选的特征组
    for group_name in excluded:
        # 获取当前要测试的组组合
        current_selected_groups = features_in_model + [group_name]
        # 拼接选中的所有特征组
        selected_features = []
        for gn in current_selected_groups:
            group_idx = group_names.index(gn)
            selected_features.append(total_features[group_idx])
        X = np.hstack(selected_features)
        
        # 训练模型并评估性能
        model = LinearRegression()
        model.fit(X, y)
        y_pred = model.predict(X)
        score = r2_score(y, y_pred)
        
        # 更新当前最优的组
        if score > current_best_score:
            current_best_score = score
            best_group = group_name
    
    # 将最优组加入模型,从排除列表移除
    features_in_model.append(best_group)
    excluded.remove(best_group)
    best_score = current_best_score
    
    print(f"加入特征组 {best_group},当前模型R²: {best_score:.4f}")

print(f"\n最终选中的特征组: {features_in_model}")

4. 向后逐步剔除的实现

如果你的需求是从全特征组开始,每次移除一个对模型性能影响最小的组,代码如下:

best_score = -np.inf

# 迭代剔除,直到只剩一个组(你可以调整停止条件)
while len(features_in_model) > 1:
    current_best_score = -np.inf
    group_to_remove = None
    
    # 遍历当前模型里的每个组,尝试移除它
    for group_name in features_in_model:
        # 临时移除当前组后的组列表
        temp_selected_groups = [gn for gn in features_in_model if gn != group_name]
        # 拼接特征矩阵
        selected_features = [total_features[group_names.index(gn)] for gn in temp_selected_groups]
        X = np.hstack(selected_features)
        
        # 评估模型性能
        model = LinearRegression()
        model.fit(X, y)
        y_pred = model.predict(X)
        score = r2_score(y, y_pred)
        
        # 找到移除后性能最好的组(也就是最该移除的组)
        if score > current_best_score:
            current_best_score = score
            group_to_remove = group_name
    
    # 移除该组,更新状态
    features_in_model.remove(group_to_remove)
    excluded.append(group_to_remove)
    best_score = current_best_score
    
    print(f"移除特征组 {group_to_remove},当前模型R²: {best_score:.4f}")

print(f"\n最终保留的特征组: {features_in_model}")

关键提示

  • 核心逻辑是跟踪特征组的选择状态,而不是直接操作拼接后的列,这样就能保证每次处理的是整个组,不会拆分单个列。
  • 如果特征组数量较多,用索引代替名称会更高效(比如用group_indices = [0,1,2],直接用索引取组,不用调用index()方法)。
  • 你可以根据自己的任务替换模型(比如Ridge、Lasso)或者评估指标(比如MSE)。

内容的提问来源于stack exchange,提问作者Shawn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:20:11