You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何编码类别数量可变的高基数分类特征?

适配Logistic Regression的高基数多类别特征编码方案

针对单索引对应多个高基数无顺序分类特征的场景,以下是几种实用的编码方法,均基于Python实现,适配Logistic Regression的线性特性:

1. 计数编码(Count Encoding)

统计每个类别在全数据集的出现频次,用频次作为编码值,既能保留类别全局分布信息,又不会造成特征维度爆炸。可对单索引下的多类别频次取均值、最大值或最小值作为最终特征(均值为常用选择)。

import pandas as pd

# 统计每个类别的全局出现次数
cat_counts = df2['f_cat'].value_counts().to_dict()
# 为df2添加频次列
df2['count'] = df2['f_cat'].map(cat_counts)
# 按idx聚合,取频次均值作为样本特征
count_encoded = df2.groupby('idx')['count'].mean().reset_index(name='f_cat_count')
# 与主数据集合并
merged_df = pd.merge(df, count_encoded, on='idx', how='left')

2. 目标编码(Target Encoding)

若有标签数据,用每个类别对应的目标变量均值作为编码值,能直接建立类别与预测目标的关联,对Logistic Regression这类线性模型效果较好。需通过交叉验证避免过拟合。

from sklearn.model_selection import KFold
import numpy as np

# 假设主数据集df包含目标列target
df_with_target = pd.merge(df, df2, on='idx', how='left')
kf = KFold(n_splits=5, shuffle=True, random_state=42)
df_with_target['target_enc'] = np.nan

for train_idx, val_idx in kf.split(df_with_target):
    train_data = df_with_target.iloc[train_idx]
    # 统计训练集中每个类别的目标均值
    cat_target_mean = train_data.groupby('f_cat')['target'].mean().to_dict()
    # 为验证集赋值
    df_with_target.loc[val_idx, 'target_enc'] = df_with_target.loc[val_idx, 'f_cat'].map(cat_target_mean)

# 按idx聚合,取均值作为样本特征
target_encoded = df_with_target.groupby('idx')['target_enc'].mean().reset_index(name='f_cat_target')
merged_df = pd.merge(df, target_encoded, on='idx', how='left')

3. 哈希编码(Hashing Encoding)

用哈希函数将高基数类别映射到固定维度的低维空间,彻底避免维度爆炸,适合类别数量极大的场景。虽可能出现哈希碰撞,但可通过调整哈希维度降低概率。

from sklearn.feature_extraction import FeatureHasher

# 将每个idx对应的类别转换为列表格式
cat_lists = df2.groupby('idx')['f_cat'].apply(list).reset_index(name='cat_list')
# 初始化哈希编码器,设置输出维度(示例设为32,可按需调整)
hasher = FeatureHasher(n_features=32, input_type='string')
# 对每个idx的类别列表进行哈希编码
hash_features = hasher.transform(cat_lists['cat_list'])
# 转换为DataFrame并与主数据集合并
hash_df = pd.DataFrame(hash_features.toarray(), index=cat_lists['idx'])
merged_df = pd.merge(df, hash_df, left_on='idx', right_index=True, how='left')

4. 嵌入编码(Embedding Encoding)

将每个类别映射为低维稠密向量,适合类别存在潜在关联的场景。可使用预训练词嵌入(若类别为文本类),或通过神经网络训练嵌入层,再将嵌入向量的均值/最大值作为特征输入Logistic Regression。

import gensim.downloader as api
import numpy as np

# 加载预训练词嵌入模型(示例用50维GloVe)
word_vectors = api.load('glove-wiki-gigaword-50')

# 计算单索引下所有类别的嵌入向量均值
def get_embedding_mean(cats):
    vecs = []
    for cat in cats:
        if cat in word_vectors:
            vecs.append(word_vectors[cat])
    return np.mean(vecs, axis=0) if vecs else np.zeros(50)

cat_lists = df2.groupby('idx')['f_cat'].apply(list).reset_index(name='cat_list')
embeddings = cat_lists['cat_list'].apply(get_embedding_mean)
# 转换为DataFrame并与主数据集合并
embedding_df = pd.DataFrame(embeddings.tolist(), index=cat_lists['idx'], columns=[f'emb_{i}' for i in range(50)])
merged_df = pd.merge(df, embedding_df, left_on='idx', right_index=True, how='left')

方法选择建议

  • 无标签数据:优先选择计数编码或哈希编码
  • 有标签数据:优先选择目标编码(务必搭配交叉验证)
  • 类别为文本且含语义关联:选择嵌入编码

内容的提问来源于stack exchange,提问作者cdeterman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 10:32:21