You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将含多分类变量的DataFrame直接转换为稀疏矩阵?

解决方案:直接从多分类变量DataFrame生成稀疏矩阵(避免稠密透视表内存溢出)

核心思路是跳过生成稠密透视表的步骤,直接基于原始数据构建稀疏矩阵——通过将每个分类特征与name列组合成唯一标识,全程仅处理非零有效特征组合,彻底避免中间稠密矩阵的内存占用问题。

方法一:底层构建CSR稀疏矩阵(内存效率最高)

import pandas as pd
from scipy.sparse import csr_matrix
import numpy as np

# 示例数据
data = {
    'id': [13,13,14,14,14,15],
    'name': ['alex', 'mary', 'alex', 'barry', 'john', 'john'],
    'categ': ['dog', 'cat', 'dog', 'ant', 'fox', 'seal'],
    'size': ['big', 'small', 'big', 'tiny', 'medium', 'big']
}
df = pd.DataFrame(data)

# 1. 生成所有有效特征标识(格式:name_特征类型_特征值)
categ_features = df.apply(lambda row: f"{row['name']}_categ_{row['categ']}", axis=1)
size_features = df.apply(lambda row: f"{row['name']}_size_{row['size']}", axis=1)
all_features = pd.concat([categ_features, size_features])
# 映射特征到唯一索引
unique_features, feature_idx = np.unique(all_features, return_inverse=True)

# 2. 映射id到唯一索引
unique_ids, id_idx = np.unique(df['id'], return_inverse=True)
# 每个id对应2个特征(categ+size),因此重复id索引
id_idx_repeated = np.repeat(id_idx, 2)

# 3. 构建CSR稀疏矩阵(非零值全为1,表示该id存在对应特征)
sparse_mat = csr_matrix(
    (np.ones(len(feature_idx), dtype=int), (id_idx_repeated, feature_idx)),
    shape=(len(unique_ids), len(unique_features))
)

# 转换为稀疏DataFrame(可选)
sparse_df = pd.DataFrame.sparse.from_spmatrix(
    sparse_mat, index=unique_ids, columns=unique_features
)
print(sparse_df.dtypes)

方法二:Pandas稀疏编码+长格式转换(代码更简洁)

import pandas as pd

# 示例数据
data = {
    'id': [13,13,14,14,14,15],
    'name': ['alex', 'mary', 'alex', 'barry', 'john', 'john'],
    'categ': ['dog', 'cat', 'dog', 'ant', 'fox', 'seal'],
    'size': ['big', 'small', 'big', 'tiny', 'medium', 'big']
}
df = pd.DataFrame(data)

# 1. 宽表转长表,统一特征格式
long_df = df.melt(
    id_vars=['id', 'name'],
    var_name='feature_type',
    value_name='feature_value'
)
long_df['full_feature'] = long_df.apply(
    lambda row: f"{row['name']}_{row['feature_type']}_{row['feature_value']}",
    axis=1
)

# 2. 稀疏编码+按id聚合
sparse_df = pd.get_dummies(long_df, columns=['full_feature'], sparse=True)\
              .groupby('id').max()

print(sparse_df.dtypes)

方案优势

  • 全程无稠密矩阵生成,内存占用仅与原始数据规模正相关,彻底解决大数据集内存溢出问题。
  • 稀疏矩阵仅存储非零有效特征组合,比稠密矩阵节省90%以上的内存空间(视数据稀疏度而定)。

内容的提问来源于stack exchange,提问作者Dudelstein

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 08:02:49