如何将scipy稀疏矩阵转为包含0值元素的pandas DataFrame
scipy稀疏矩阵csr_matrix转含0值的pandas DataFrame方案
csr矩阵转coo格式后仅存储非零元素,因此直接提取属性生成DataFrame只会得到非零值对应的条目,如需包含所有0值条目可以参考以下两种方案:
方案1:适用于尺寸较小的矩阵
直接将稀疏矩阵转为稠密矩阵后处理,代码简洁易读:
import numpy as np import pandas as pd from scipy.sparse import csr_matrix csr = csr_matrix(np.array( [[0, 0, 4], [1, 0, 0], [2, 0, 0],])) # 转为稠密矩阵后堆叠行列,重置索引即可得到全量条目 df = pd.DataFrame(csr.todense())\ .stack()\ .reset_index() # 重命名为目标列名 df.columns = ['index', 'col', 'data']
方案2:适用于大尺寸稀疏矩阵
避免转为稠密矩阵占用过多内存,通过全量行列索引重映射实现:
import numpy as np import pandas as pd from scipy.sparse import csr_matrix csr = csr_matrix(np.array( [[0, 0, 4], [1, 0, 0], [2, 0, 0],])) coo = csr.tocoo(copy=False) n_rows, n_cols = csr.shape # 生成所有行、列的全量组合索引 full_midx = pd.MultiIndex.from_product([np.arange(n_rows), np.arange(n_cols)], names=['index', 'col']) # 基于非零值构建Series后重索引,缺失值填充为0 data_series = pd.Series(coo.data, index=pd.MultiIndex.from_arrays([coo.row, coo.col], names=['index', 'col'])) df = data_series.reindex(full_midx, fill_value=0).reset_index(name='data')
两种方案输出结果都符合预期:
index col data 0 0 0 0 1 0 1 0 2 0 2 4 3 1 0 1 4 1 1 0 5 1 2 0 6 2 0 2 7 2 1 0 8 2 2 0
内容的提问来源于stack exchange,提问作者Chen
相关产品推荐
相关产品推荐

