基于Numpy特征组的回归逐步子集选择器构建疑问
回归逐步子集选择器:特征组层面的筛选实现
我明白你现在的需求——之前能搞定单特征列的逐步选择,但现在要升级到按特征组(比如a、b、c这种整组)来做逐步筛选,对吧?其实核心思路就是把操作单元从「单个特征列」换成「整个特征组」,不用直接去碰拼接后的大矩阵,而是先跟踪哪些组在模型里、哪些被排除,每次迭代基于选中的组拼接特征再评估。下面给你具体的实现方案:
1. 先给特征组做标识
首先给每个特征组起个好区分的名字或者用索引,方便我们跟踪状态。比如用名称的话:
import numpy as np # 你的原始特征组 a = np.array([[ 1., 1.], [ 1., 1.], [ 1., 1.]]) b = np.array([[ 88., 42.5, 9. ], [ 121.5, 76., 42.5], [ 167., 121.5, 88. ]]) c = np.array([[ 88., 42.5, 13. ], [ 117.5, 72., 42.5], [ 163., 117.5, 88. ]]) # 定义特征组的名称和对应的组列表 group_names = ['a', 'b', 'c'] total_features = [a, b, c]
2. 初始化特征组的筛选状态
如果是向前逐步选择(从无到有,每次加最优组),初始状态是模型里没有任何组,所有组都在排除列表里:
features_in_model = [] # 当前模型包含的特征组 excluded = group_names.copy() # 初始时所有组都待选
如果是向后逐步剔除(从全组开始,每次删性能影响最小的组),初始状态是模型包含所有组,排除列表为空:
features_in_model = group_names.copy() # 初始包含所有组 excluded = [] # 初始无被排除的组
3. 完整的向前逐步选择实现
下面是带模型评估的完整代码,用R²作为性能指标(你可以换成MSE、MAE等):
from sklearn.linear_model import LinearRegression from sklearn.metrics import r2_score # 模拟目标变量(你替换成自己的真实y即可) y = np.array([10., 20., 30.]) best_score = -np.inf # 迭代筛选,直到所有组都加入模型 while excluded: current_best_score = -np.inf best_group = None # 遍历所有待选的特征组 for group_name in excluded: # 获取当前要测试的组组合 current_selected_groups = features_in_model + [group_name] # 拼接选中的所有特征组 selected_features = [] for gn in current_selected_groups: group_idx = group_names.index(gn) selected_features.append(total_features[group_idx]) X = np.hstack(selected_features) # 训练模型并评估性能 model = LinearRegression() model.fit(X, y) y_pred = model.predict(X) score = r2_score(y, y_pred) # 更新当前最优的组 if score > current_best_score: current_best_score = score best_group = group_name # 将最优组加入模型,从排除列表移除 features_in_model.append(best_group) excluded.remove(best_group) best_score = current_best_score print(f"加入特征组 {best_group},当前模型R²: {best_score:.4f}") print(f"\n最终选中的特征组: {features_in_model}")
4. 向后逐步剔除的实现
如果你的需求是从全特征组开始,每次移除一个对模型性能影响最小的组,代码如下:
best_score = -np.inf # 迭代剔除,直到只剩一个组(你可以调整停止条件) while len(features_in_model) > 1: current_best_score = -np.inf group_to_remove = None # 遍历当前模型里的每个组,尝试移除它 for group_name in features_in_model: # 临时移除当前组后的组列表 temp_selected_groups = [gn for gn in features_in_model if gn != group_name] # 拼接特征矩阵 selected_features = [total_features[group_names.index(gn)] for gn in temp_selected_groups] X = np.hstack(selected_features) # 评估模型性能 model = LinearRegression() model.fit(X, y) y_pred = model.predict(X) score = r2_score(y, y_pred) # 找到移除后性能最好的组(也就是最该移除的组) if score > current_best_score: current_best_score = score group_to_remove = group_name # 移除该组,更新状态 features_in_model.remove(group_to_remove) excluded.append(group_to_remove) best_score = current_best_score print(f"移除特征组 {group_to_remove},当前模型R²: {best_score:.4f}") print(f"\n最终保留的特征组: {features_in_model}")
关键提示
- 核心逻辑是跟踪特征组的选择状态,而不是直接操作拼接后的列,这样就能保证每次处理的是整个组,不会拆分单个列。
- 如果特征组数量较多,用索引代替名称会更高效(比如用
group_indices = [0,1,2],直接用索引取组,不用调用index()方法)。 - 你可以根据自己的任务替换模型(比如Ridge、Lasso)或者评估指标(比如MSE)。
内容的提问来源于stack exchange,提问作者Shawn
相关产品推荐
相关产品推荐

