使用scipy.sparse.bmat构建超大稀疏矩阵出现negative row index错误
问题根源分析
你推测的索引溢出是正确的,你执行的scipy.sparse.sputils.get_index_dtype测试仅验证了工具函数的标称能力,没有覆盖scipy.sparse.bmat的内部实现缺陷:
- Python3.7对应的scipy 1.5.x及更早版本中,
bmat拼接子矩阵计算偏移量时,默认使用int32类型存储临时索引,当总行列数超过2^31-1(约21.47亿)时,正整数溢出会被解释为负数,直接触发negative row index found报错 - 即使numpy全局支持int64索引,该版本bmat的内部偏移量计算逻辑也没有主动指定dtype为int64,会发生32位溢出截断
可行的解决方法
- 方案1:升级scipy到1.7.0及以上版本,该版本已经修复了bmat大索引溢出问题,会主动给索引数组指定int64的dtype
- 方案2:如果无法升级scipy,手动实现拼接逻辑,避免调用bmat,示例代码如下:
import numpy as np from scipy.sparse import csr_matrix # 这里替换为你的实际总行列数 total_rows = 2600000000 total_cols = 2600000000 all_data = [] all_row = [] all_col = [] row_block_cnt = len(HList) col_block_cnt = len(HList[0]) current_row_offset = 0 for i in range(row_block_cnt): current_col_offset = 0 for j in range(col_block_cnt): sub_mat = HList[i][j] # 跳过空块 if sub_mat is None or sub_mat.nnz == 0: current_col_offset += HList[0][j].shape[1] continue sub_coo = sub_mat.tocoo() all_data.append(sub_coo.data) # 强制转int64后再加偏移,从根源避免溢出 all_row.append(sub_coo.row.astype(np.int64) + current_row_offset) all_col.append(sub_coo.col.astype(np.int64) + current_col_offset) current_col_offset += sub_coo.shape[1] current_row_offset += HList[i][0].shape[0] # 拼接所有数组生成最终矩阵 all_data = np.concatenate(all_data) all_row = np.concatenate(all_row) all_col = np.concatenate(all_col) final_mat = csr_matrix((all_data, (all_row, all_col)), shape=(total_rows, total_cols))
- 方案3:如果不想重写拼接逻辑,可以临时修改本地的
scipy/sparse/construct.py中bmat函数的实现,在生成row、col数组时强制指定dtype=np.int64
内容的提问来源于stack exchange,提问作者Theo
相关产品推荐
相关产品推荐

