读取保存后的npz文件创建CSR矩阵时出现OOM问题求助
问题:npy转CSR矩阵正常,转存npz后再转CSR出现OOM错误
我有一份npy格式存储的稀疏矩阵数据,直接读取并转换为CSR格式时完全正常,但将这份npy文件读取后保存为npz文件,再从npz读取数据创建CSR矩阵时,出现了Detected 1 oom-kill event(s)内存不足错误。两次创建CSR的行列数和数据量完全一致,但OOM问题依然发生,无法定位根源。
直接读取npy转CSR的正常代码
def load_adjacency_matrix_csr(folder: str, filename: str, suffix: str = "npy", row_idx: int = 1, num_rows: int = None, dtype: np.dtype = np.float32) -> sparse.csr_matrix: coo_indices = np.load(os.path.join(folder, f"{filename}.{suffix}")) rows = coo_indices[row_idx] cols = coo_indices[1 - row_idx] data = np.ones(len(rows), dtype=dtype) num_rows = num_rows or rows.max() + 1 if num_rows < rows.max() + 1: raise ValueError("The number of rows in the file is larger than the specified number of rows.") csr_mat = sparse.csr_matrix((data, (rows, cols)), shape=(num_rows, num_rows), dtype=dtype) return csr_mat
将npy转存为npz的代码
def save_coo_matrix(npx: int, folder: str, filename: str, suffix: str = "npy", row_idx: int = 1, dtype: np.dtype = np.float32) -> None: file_path = os.path.join(folder, f"{filename}.{suffix}") coo_indices = np.load(file_path) rows = coo_indices[row_idx] cols = coo_indices[1 - row_idx] filename += "_temp" file_path = os.path.join(folder, f"{filename}.npz") np.savez(file_path, rows=rows, cols=cols)
读取npz转CSR报错的代码
def load_adjacency_matrix_csr(npx: int, folder: str, filename: str, suffix: str = "npy", row_idx: int = 1, num_rows: int = None, dtype: np.dtype = np.float32) -> sparse.csr_matrix: file_path = os.path.join(folder, f"{filename}_padded.npz") with np.load(file_path) as data: rows = data["rows"] cols = data["cols"] num_rows = max(rows) + 1 num_cols = max(cols) + 1 values = np.ones(len(rows)) # fails here when creating the csr matrix (the num_rows is the same value as before) csr_mat = sparse.csr_matrix((values, (rows, cols)), shape=(num_rows, num_cols)) return csr_mat
问题根源分析
- 数据类型不一致:原代码中
data指定了dtype=np.float32,但新代码中values = np.ones(len(rows))默认生成np.float64类型,内存占用直接翻倍。 - 整数类型膨胀:
np.savez可能会将原本紧凑的整数类型(如np.int32)转换为np.int64存储,加载后数组内存占用增加。 - 形状定义差异:原代码强制使用方阵
shape=(num_rows, num_rows),新代码用(num_rows, num_cols),即使数值相同,scipy创建CSR时的临时内存分配逻辑可能产生额外开销。 - 手动拆分存储的冗余:手动拆分
rows/cols保存,不如scipy原生稀疏矩阵存储高效,加载时可能产生临时数组占用内存。
解决方案
- 统一数据类型:创建
values时显式指定与原代码一致的dtype:values = np.ones(len(rows), dtype=dtype) - 强制紧凑整数类型:保存和加载时将行列索引转换为内存占用更小的整数类型:
在save_coo_matrix中:
在加载时:rows = coo_indices[row_idx].astype(np.int32) cols = coo_indices[1 - row_idx].astype(np.int32)rows = data["rows"].astype(np.int32) cols = data["cols"].astype(np.int32) - 保持方阵形状定义:和原代码一致使用方阵形状(如果矩阵确实是方阵):
csr_mat = sparse.csr_matrix((values, (rows, cols)), shape=(num_rows, num_rows), dtype=dtype) - 使用scipy原生稀疏矩阵存储:跳过手动拆分,直接保存/加载稀疏矩阵,避免格式转换问题:
# 保存稀疏矩阵 def save_sparse_matrix(folder: str, filename: str, mat: sparse.csr_matrix): sparse.save_npz(os.path.join(folder, f"{filename}.npz"), mat) # 加载稀疏矩阵 def load_sparse_matrix(folder: str, filename: str) -> sparse.csr_matrix: return sparse.load_npz(os.path.join(folder, f"{filename}.npz"))
内容的提问来源于stack exchange,提问作者Andrew Mathews
相关产品推荐
相关产品推荐

