You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何Numpy数组比等价Pickle序列化列表大?求压缩方案

问题解析与解决方案

为什么Numpy数组比Pickle序列化列表大?

1. Numpy的存储特性

你的10000×10000×7尺寸、dtype=np.float16数组,理论体积计算为:
10000 * 10000 * 7 * 2字节 = 1.4GB(每个float16占2字节)。np.save默认保存未压缩的原始二进制数据,不管数组内有多少重复的子数组,都会平铺存储所有元素,所以实际文件大小和理论值完全一致。

2. Pickle的自动优化

Pickle序列化Python列表时,会对重复对象做引用优化:你用random.choice采样时,大量子列表是从原列表重复选取的,Pickle只会把相同的子列表存储一次,后续重复出现的只保存引用指针,避免了重复数据的冗余存储,因此文件体积大幅缩小到约500MB。

如何缩小Numpy数组的保存体积?

方法一:启用Numpy的压缩存储

用np.savez_compressed替代np.save,它会通过zlib算法压缩数据,针对有大量重复值的数组,压缩率会非常高,能显著降低文件体积。示例代码:

def sample_games(all_games, file_name):
    all_games = np.array(all_games, dtype=np.float16)
    DRAW = 10000
    SAMPLE = 10000
    sampled_indices = rng.choice(all_games.shape[0], size=(SAMPLE, DRAW), replace=True)
    sampled_data = all_games[sampled_indices]

    np.savez_compressed(file_name, sampled_data=sampled_data)

读取时通过np.load(file_name)['sampled_data']还原数组。

方法二:去重后存储索引+唯一数据

利用采样数据的重复性,先提取唯一子数组,再存储唯一数组和对应的索引,能极大节省空间:

def sample_games(all_games, file_name):
    all_games = np.array(all_games, dtype=np.float16)
    DRAW = 10000
    SAMPLE = 10000
    sampled_indices = rng.choice(all_games.shape[0], size=(SAMPLE, DRAW), replace=True)
    sampled_data = all_games[sampled_indices]

    # 提取唯一子数组和索引
    unique_games, indices = np.unique(sampled_data.reshape(-1,7), axis=0, return_inverse=True)
    # 保存压缩后的唯一数组和索引
    np.savez_compressed(file_name, unique_games=unique_games, indices=indices.reshape(SAMPLE, DRAW))

读取还原代码:

loaded_data = np.load(file_name)
sampled_data_restored = loaded_data['unique_games'][loaded_data['indices']]

方法三:使用高效压缩存储库

对于大型重复数组,zarr或h5py这类支持分块压缩的库会比Numpy原生方法更高效,它们能针对数组的重复区域做更精细的压缩优化。

内容的提问来源于stack exchange,提问作者findingmyway

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 05:45:24