You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中基于特征重要性的循环特征选择实现(保留≥150特征)

基于XGBoost的迭代式特征选择实现方法

需求说明

现有包含二分类目标变量Y和预测变量X1~Xn的Pandas DataFrame,需将原两步手动特征选择逻辑整合为单一循环:

  • 每次迭代优先删除特征重要性为0的特征
  • 若不存在重要性为0的特征,则删除重要性低于当前均值的特征
  • 最终保留至少150个特征

实现代码与说明

核心函数实现

import pandas as pd
from xgboost import XGBClassifier

def iterative_feature_selection(X, y, min_features=150):
    current_features = X.columns.tolist()
    
    while len(current_features) > min_features:
        # 训练XGBoost模型,设置参数避免警告并保证可复现
        xgb_model = XGBClassifier(
            random_state=42,
            use_label_encoder=False,
            eval_metric='logloss'
        )
        xgb_model.fit(X[current_features], y)
        
        # 提取当前特征的重要性
        feature_importance = pd.Series(
            xgb_model.feature_importances_,
            index=current_features
        )
        
        # 优先处理重要性为0的特征
        zero_importance = feature_importance[feature_importance == 0].index
        if not zero_importance.empty:
            current_features = [f for f in current_features if f not in zero_importance]
            print(f"移除{len(zero_importance)}个0重要性特征,剩余{len(current_features)}个")
        else:
            # 计算当前重要性均值,筛选低于均值的特征
            importance_mean = feature_importance.mean()
            below_mean = feature_importance[feature_importance < importance_mean].index
            
            # 避免删除后特征数低于最小值,只删除到刚好满足要求
            if len(current_features) - len(below_mean) < min_features:
                need_remove = len(current_features) - min_features
                # 按重要性升序取最不重要的need_remove个
                below_mean = feature_importance.sort_values().index[:need_remove]
            
            current_features = [f for f in current_features if f not in below_mean]
            print(f"移除{len(below_mean)}个低重要性特征,剩余{len(current_features)}个")
    
    return current_features, X[current_features]

关键逻辑说明

  1. 循环终止条件:当剩余特征数≤150时停止迭代
  2. 动态更新重要性:每次迭代都重新训练XGBoost,确保特征集合变化后,重要性计算是基于当前特征的
  3. 优先级处理:严格遵循“先删0重要性特征,再删低于均值特征”的规则
  4. 边界保护:当删除低于均值的特征会导致总数不足150时,仅删除多余的部分,保证最终至少保留150个特征
  5. 参数配置:设置random_state保证结果可复现,关闭use_label_encoder并指定eval_metric避免XGBoost版本兼容警告

使用示例

# 假设你的数据集存储在df中,Y是目标列
X = df.drop('Y', axis=1)
y = df['Y']

# 执行特征选择
final_features, final_X = iterative_feature_selection(X, y, min_features=150)

内容的提问来源于stack exchange,提问作者dingaro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 20:05:30