如何解决循环中scipy csc_matrix索引指针需从0开始的警告?
问题:分批加载Julia保存的CSC稀疏矩阵时Scipy报错"index pointer should start with 0"
我们分析大型数据集时,在Julia中保存为CSC稀疏数组,之后用Python的scipy.csc_matrix加载时因内存不足(提示无法分配93GiB),尝试循环分批处理数据,但运行时报错"index pointer should start with 0",第二次迭代就终止。
原代码
def get_jld(filename): orient_list=[] ID_num=999 f = h5py.File(filename, 'r') data= f["sm"][()] column_ptr=f[data[2]][:]-1 ## 修正Julia的1-based索引为0-based indices=f[data[3]][:]-1 ## 修正索引 values =f[data[4]][:] for i in range(0, data[1], ID_num): out = pd.DataFrame(csc_matrix((values,indices,column_ptr[i:i+ID_num+1]), shape=(data[0],ID_num)).toarray()) orient_angle=find_orientation(out) orient_list.append(orient_angle) f.close() return orient_list
报错信息
ValueError Traceback (most recent call last) Cell In[6], line 2 1 filename=vid_list[0] ----> 2 test = get_jld(filename) Cell In[3], line 13, in get_jld(filename) 10 #indices[indices < time_start] = time_start 11 #print([data[0],data[1]]) 12 for i in range(0, data[1], ID_num): ---> 13 out =pd.DataFrame(csc_matrix((values,indices,column_ptr[i:i+ID_num+1]), shape=(data[0],ID_num)).toarray()) 14 orient_angle=find_orientation(out) 15 orient_list.append(orient_angle) File ~\anaconda3\envs\qlm_analysis\lib\site-packages\scipy\sparse\_compressed.py:106, in _cs_matrix.__init__(self, arg1, shape, dtype, copy) 103 if dtype is not None: 104 self.data = self.data.astype(dtype, copy=False) --> 106 self.check_format(full_check=False) File ~\anaconda3\envs\qlm_analysis\lib\site-packages\scipy\sparse\_compressed.py:172, in _cs_matrix.check_format(self, full_check) 169 raise ValueError("index pointer size ({}) should be ({})" 170 "".format(len(self.indptr), major_dim + 1)) 171 if (self.indptr[0] != 0): --> 172 raise ValueError("index pointer should start with 0") 174 # check index and data arrays 175 if (len(self.indices) != len(self.data)): ValueError: index pointer should start with 0
解决方案
报错原因是:每一批构造独立CSC矩阵时,indptr(即代码中的column_ptr切片)必须以0开头,但原代码第二次迭代时,切片的第一个元素是原数组的column_ptr[i],并非0。
修改思路:
- 对每一批,提取对应范围内的
values和indices子数组 - 重新构造该批次的
indptr,使其从0开始(通过减去该批次的起始指针值) - 处理最后一批的边界情况,避免越界
修改后的代码:
import pandas as pd from scipy.sparse import csc_matrix import h5py def get_jld(filename): orient_list = [] ID_num = 999 f = h5py.File(filename, 'r') data = f["sm"][()] column_ptr = f[data[2]][:] - 1 # 修正Julia的1-based索引为0-based indices = f[data[3]][:] - 1 # 修正索引 values = f[data[4]][:] total_cols = data[1] for i in range(0, total_cols, ID_num): # 确定当前批次的列范围,处理最后一批的边界 end_col = min(i + ID_num, total_cols) batch_cols = end_col - i # 获取当前批次对应的指针范围 start_ptr = column_ptr[i] end_ptr = column_ptr[end_col] # 提取当前批次的非零元素和对应索引 values_sub = values[start_ptr:end_ptr] indices_sub = indices[start_ptr:end_ptr] # 构造当前批次的indptr,确保以0开头 indptr_sub = column_ptr[i:end_col+1] - start_ptr # 构造批次CSC矩阵并转换为DataFrame csc_batch = csc_matrix((values_sub, indices_sub, indptr_sub), shape=(data[0], batch_cols)) out = pd.DataFrame(csc_batch.toarray()) orient_angle = find_orientation(out) orient_list.append(orient_angle) f.close() return orient_list
关键修改点
- 计算每一批的实际结束列,避免最后一批越界
- 只提取当前批次的非零元素,减少内存占用
- 调整批次
indptr的起始值为0,符合Scipy的CSC格式要求
内容的提问来源于stack exchange,提问作者cylee
相关产品推荐
相关产品推荐

