为何Numpy数组比等价Pickle序列化列表大?求压缩方案
问题解析与解决方案
为什么Numpy数组比Pickle序列化列表大?
1. Numpy的存储特性
你的10000×10000×7尺寸、dtype=np.float16数组,理论体积计算为:10000 * 10000 * 7 * 2字节 = 1.4GB(每个float16占2字节)。np.save默认保存未压缩的原始二进制数据,不管数组内有多少重复的子数组,都会平铺存储所有元素,所以实际文件大小和理论值完全一致。
2. Pickle的自动优化
Pickle序列化Python列表时,会对重复对象做引用优化:你用random.choice采样时,大量子列表是从原列表重复选取的,Pickle只会把相同的子列表存储一次,后续重复出现的只保存引用指针,避免了重复数据的冗余存储,因此文件体积大幅缩小到约500MB。
如何缩小Numpy数组的保存体积?
方法一:启用Numpy的压缩存储
用np.savez_compressed替代np.save,它会通过zlib算法压缩数据,针对有大量重复值的数组,压缩率会非常高,能显著降低文件体积。示例代码:
def sample_games(all_games, file_name): all_games = np.array(all_games, dtype=np.float16) DRAW = 10000 SAMPLE = 10000 sampled_indices = rng.choice(all_games.shape[0], size=(SAMPLE, DRAW), replace=True) sampled_data = all_games[sampled_indices] np.savez_compressed(file_name, sampled_data=sampled_data)
读取时通过np.load(file_name)['sampled_data']还原数组。
方法二:去重后存储索引+唯一数据
利用采样数据的重复性,先提取唯一子数组,再存储唯一数组和对应的索引,能极大节省空间:
def sample_games(all_games, file_name): all_games = np.array(all_games, dtype=np.float16) DRAW = 10000 SAMPLE = 10000 sampled_indices = rng.choice(all_games.shape[0], size=(SAMPLE, DRAW), replace=True) sampled_data = all_games[sampled_indices] # 提取唯一子数组和索引 unique_games, indices = np.unique(sampled_data.reshape(-1,7), axis=0, return_inverse=True) # 保存压缩后的唯一数组和索引 np.savez_compressed(file_name, unique_games=unique_games, indices=indices.reshape(SAMPLE, DRAW))
读取还原代码:
loaded_data = np.load(file_name) sampled_data_restored = loaded_data['unique_games'][loaded_data['indices']]
方法三:使用高效压缩存储库
对于大型重复数组,zarr或h5py这类支持分块压缩的库会比Numpy原生方法更高效,它们能针对数组的重复区域做更精细的压缩优化。
内容的提问来源于stack exchange,提问作者findingmyway
相关产品推荐
相关产品推荐

