You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用scipy.sparse.bmat构建超大稀疏矩阵出现negative row index错误

问题根源分析

你推测的索引溢出是正确的,你执行的scipy.sparse.sputils.get_index_dtype测试仅验证了工具函数的标称能力,没有覆盖scipy.sparse.bmat的内部实现缺陷:

  • Python3.7对应的scipy 1.5.x及更早版本中,bmat拼接子矩阵计算偏移量时,默认使用int32类型存储临时索引,当总行列数超过2^31-1(约21.47亿)时,正整数溢出会被解释为负数,直接触发negative row index found报错
  • 即使numpy全局支持int64索引,该版本bmat的内部偏移量计算逻辑也没有主动指定dtype为int64,会发生32位溢出截断
可行的解决方法
  • 方案1:升级scipy到1.7.0及以上版本,该版本已经修复了bmat大索引溢出问题,会主动给索引数组指定int64的dtype
  • 方案2:如果无法升级scipy,手动实现拼接逻辑,避免调用bmat,示例代码如下:
import numpy as np
from scipy.sparse import csr_matrix

# 这里替换为你的实际总行列数
total_rows = 2600000000
total_cols = 2600000000

all_data = []
all_row = []
all_col = []

row_block_cnt = len(HList)
col_block_cnt = len(HList[0])

current_row_offset = 0
for i in range(row_block_cnt):
    current_col_offset = 0
    for j in range(col_block_cnt):
        sub_mat = HList[i][j]
        # 跳过空块
        if sub_mat is None or sub_mat.nnz == 0:
            current_col_offset += HList[0][j].shape[1]
            continue
        sub_coo = sub_mat.tocoo()
        all_data.append(sub_coo.data)
        # 强制转int64后再加偏移,从根源避免溢出
        all_row.append(sub_coo.row.astype(np.int64) + current_row_offset)
        all_col.append(sub_coo.col.astype(np.int64) + current_col_offset)
        current_col_offset += sub_coo.shape[1]
    current_row_offset += HList[i][0].shape[0]

# 拼接所有数组生成最终矩阵
all_data = np.concatenate(all_data)
all_row = np.concatenate(all_row)
all_col = np.concatenate(all_col)
final_mat = csr_matrix((all_data, (all_row, all_col)), shape=(total_rows, total_cols))
  • 方案3:如果不想重写拼接逻辑,可以临时修改本地的scipy/sparse/construct.py中bmat函数的实现,在生成row、col数组时强制指定dtype=np.int64

内容的提问来源于stack exchange,提问作者Theo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 13:54:04