使用VarianceThreshold()识别准常分类变量未达预期,如何修正?
解决分类变量准常量过滤问题
问题根源
你用OrdinalEncoder编码分类变量后,VarianceThreshold计算的是数值方差,而非分类变量的类别占比方差。比如Alley特征94%取值为2,剩余6%为0/1,数值方差会高于你设置的threshold=0.1,导致它没被识别为准常量特征。
两种可行修改方案
方案1:直接按类别占比过滤(推荐,直观高效)
跳过编码步骤,直接计算每个分类特征的最高类别占比,过滤掉占比超过阈值的特征:
import pandas as pd def remove_quasi_constant_cats(df, threshold=0.9): quasi_constant_cols = [] # 遍历所有分类特征 for col in df.select_dtypes(include=['object', 'category']).columns: # 计算最高类别占比 top_ratio = df[col].value_counts(normalize=True).iloc[0] if top_ratio >= threshold: quasi_constant_cols.append(col) # 返回清洗后的数据集和被移除的特征列表 return df.drop(quasi_constant_cols, axis=1), quasi_constant_cols # 使用示例:填充空值后调用函数 df_filled = df.fillna("null_val") df_cleaned, removed_cols = remove_quasi_constant_cats(df_filled, threshold=0.94) # Alley会被包含在removed_cols中
方案2:用二分类变量配合VarianceThreshold
如果一定要用VarianceThreshold,先把每个分类特征转为「是否为最频繁类别」的二分类变量,再计算方差过滤:
from sklearn.feature_selection import VarianceThreshold import pandas as pd def remove_quasi_constant_with_var(df, threshold=0.1): df_filled = df.fillna("null_val") binary_df = pd.DataFrame() col_map = {} # 为每个分类特征生成二分类指示变量 for col in df_filled.select_dtypes(include=['object', 'category']).columns: top_val = df_filled[col].value_counts(normalize=True).index[0] binary_col = f"{col}_is_top" binary_df[binary_col] = (df_filled[col] == top_val).astype(int) col_map[binary_col] = col # 用VarianceThreshold筛选低方差特征 selector = VarianceThreshold(threshold=threshold) selector.fit(binary_df) # 映射回原特征 removed_binary_cols = binary_df.columns[~selector.get_support()] removed_cols = [col_map[col] for col in removed_binary_cols] return df_filled.drop(removed_cols, axis=1), removed_cols # 使用示例 df_cleaned, removed_cols = remove_quasi_constant_with_var(df, threshold=0.1) # Alley对应的二分类变量方差为0.94*0.06=0.0564 < 0.1,会被过滤
内容的提问来源于stack exchange,提问作者Álvaro V.
相关产品推荐
相关产品推荐

