使用np.save()保存大文件触发MemoryError,求正确保存方法
解决大数组保存时的MemoryError问题
当用np.save()或pickle.dump()保存超大数组时,会因为一次性将整个数据序列化到内存导致内存不足,以下是几种可行的解决方案:
1. 分块保存与拼接
把大数组拆分成多个小块分别存储,加载时再拼接回原数组:
import numpy as np # 分块保存 large_array = ... # 你的目标大数组 # 根据内存情况调整分块数量 chunk_count = 10 chunks = np.array_split(large_array, chunk_count) for idx, chunk in enumerate(chunks): np.save(f"large_array_chunk_{idx}.npy", chunk) # 加载并拼接 loaded_chunks = [] for idx in range(chunk_count): chunk = np.load(f"large_array_chunk_{idx}.npy") loaded_chunks.append(chunk) full_array = np.concatenate(loaded_chunks)
2. 使用内存映射文件(np.memmap)
直接在磁盘上创建映射文件,操作时仅将需要的部分加载到内存,避免占用大量内存:
import numpy as np # 写入内存映射文件 large_array = ... shape = large_array.shape dtype = large_array.dtype # 创建可读写的内存映射文件 memmap_obj = np.memmap("large_array_memmap.npy", dtype=dtype, mode='w+', shape=shape) # 将数据写入映射文件 memmap_obj[:] = large_array[:] # 释放对象确保数据写入磁盘 del memmap_obj # 加载内存映射文件 loaded_memmap = np.memmap("large_array_memmap.npy", dtype=dtype, mode='r', shape=shape) # 可像普通数组一样切片操作,仅加载对应部分到内存 partial_data = loaded_memmap[1000:2000]
3. 使用HDF5格式存储(h5py库)
HDF5支持分块存储和数据压缩,专为超大数据集设计,需要先安装h5py:
pip install h5py
保存与加载示例:
import h5py import numpy as np # 保存数据 with h5py.File("large_array.h5", "w") as h5_file: # chunks=True自动分块,compression可选开启压缩节省磁盘空间 h5_file.create_dataset("dataset_name", data=large_array, chunks=True, compression="gzip") # 加载数据 with h5py.File("large_array.h5", "r") as h5_file: # 按需加载,比如只加载前1000行:h5_file["dataset_name"][:1000] full_array = h5_file["dataset_name"][:]
内容的提问来源于stack exchange,提问作者daaa
相关产品推荐
相关产品推荐

