You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决循环中scipy csc_matrix索引指针需从0开始的警告?

问题:分批加载Julia保存的CSC稀疏矩阵时Scipy报错"index pointer should start with 0"

我们分析大型数据集时,在Julia中保存为CSC稀疏数组,之后用Python的scipy.csc_matrix加载时因内存不足(提示无法分配93GiB),尝试循环分批处理数据,但运行时报错"index pointer should start with 0",第二次迭代就终止。

原代码

def get_jld(filename):
    orient_list=[]
    ID_num=999
    f = h5py.File(filename, 'r')
    data= f["sm"][()]

    column_ptr=f[data[2]][:]-1 ## 修正Julia的1-based索引为0-based
    indices=f[data[3]][:]-1 ## 修正索引
    values =f[data[4]][:]
    for i in range(0, data[1], ID_num):
        out = pd.DataFrame(csc_matrix((values,indices,column_ptr[i:i+ID_num+1]), shape=(data[0],ID_num)).toarray())
        orient_angle=find_orientation(out)
        orient_list.append(orient_angle)
    f.close()
    return orient_list

报错信息

ValueError                                Traceback (most recent call last)

Cell In[6], line 2
      1 filename=vid_list[0]
----> 2 test = get_jld(filename)

Cell In[3], line 13, in get_jld(filename)
     10 #indices[indices < time_start] = time_start
     11 #print([data[0],data[1]])
     12 for i in range(0, data[1], ID_num):
---> 13     out =pd.DataFrame(csc_matrix((values,indices,column_ptr[i:i+ID_num+1]), shape=(data[0],ID_num)).toarray())
     14     orient_angle=find_orientation(out)
     15     orient_list.append(orient_angle)

File ~\anaconda3\envs\qlm_analysis\lib\site-packages\scipy\sparse\_compressed.py:106, in _cs_matrix.__init__(self, arg1, shape, dtype, copy)
    103 if dtype is not None:
    104     self.data = self.data.astype(dtype, copy=False)
--> 106 self.check_format(full_check=False)

File ~\anaconda3\envs\qlm_analysis\lib\site-packages\scipy\sparse\_compressed.py:172, in _cs_matrix.check_format(self, full_check)
    169     raise ValueError("index pointer size ({}) should be ({})"
    170                      "".format(len(self.indptr), major_dim + 1))
    171 if (self.indptr[0] != 0):
--> 172     raise ValueError("index pointer should start with 0")
    174 # check index and data arrays
    175 if (len(self.indices) != len(self.data)):

ValueError: index pointer should start with 0

解决方案

报错原因是:每一批构造独立CSC矩阵时,indptr(即代码中的column_ptr切片)必须以0开头,但原代码第二次迭代时,切片的第一个元素是原数组的column_ptr[i],并非0。

修改思路:

  • 对每一批,提取对应范围内的values和indices子数组
  • 重新构造该批次的indptr,使其从0开始(通过减去该批次的起始指针值)
  • 处理最后一批的边界情况,避免越界

修改后的代码:

import pandas as pd
from scipy.sparse import csc_matrix
import h5py

def get_jld(filename):
    orient_list = []
    ID_num = 999
    f = h5py.File(filename, 'r')
    data = f["sm"][()]

    column_ptr = f[data[2]][:] - 1  # 修正Julia的1-based索引为0-based
    indices = f[data[3]][:] - 1     # 修正索引
    values = f[data[4]][:]
    total_cols = data[1]

    for i in range(0, total_cols, ID_num):
        # 确定当前批次的列范围,处理最后一批的边界
        end_col = min(i + ID_num, total_cols)
        batch_cols = end_col - i

        # 获取当前批次对应的指针范围
        start_ptr = column_ptr[i]
        end_ptr = column_ptr[end_col]

        # 提取当前批次的非零元素和对应索引
        values_sub = values[start_ptr:end_ptr]
        indices_sub = indices[start_ptr:end_ptr]

        # 构造当前批次的indptr,确保以0开头
        indptr_sub = column_ptr[i:end_col+1] - start_ptr

        # 构造批次CSC矩阵并转换为DataFrame
        csc_batch = csc_matrix((values_sub, indices_sub, indptr_sub), shape=(data[0], batch_cols))
        out = pd.DataFrame(csc_batch.toarray())

        orient_angle = find_orientation(out)
        orient_list.append(orient_angle)
    
    f.close()
    return orient_list

关键修改点

  • 计算每一批的实际结束列,避免最后一批越界
  • 只提取当前批次的非零元素,减少内存占用
  • 调整批次indptr的起始值为0,符合Scipy的CSC格式要求

内容的提问来源于stack exchange,提问作者cylee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 21:44:54