You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

大文件写入提速:显微镜扫描数据存储方案优化及代码修正

显微镜扫描数据写入优化方案(兼容原有格式)

核心优化思路

原有pickle写入大数组时,因序列化开销导致速度骤降。改用struct固定格式存储配置参数 + 直接二进制写入numpy数组的方式,将IO操作从序列化大内存块转为原生二进制写入,速度可提升数倍。同时通过格式适配,保证原有读取逻辑(无论是pickle加载还是自定义读取)都能兼容。

修正后的Struct实现代码

写入代码

import struct
import numpy as np
import pickle

def save_data(file_path, xy_step, x_range, y_range, data):
    # 定义配置参数的struct格式(小端字节序,与numpy、pickle默认存储字节序一致)
    # 格式说明:
    # <: 小端字节序
    # f: float32(4字节),对应x_range[0], x_range[1], y_range[0], y_range[1], xy_step
    # i: int32(4字节),对应数组的行数、列数
    # 10s: 固定10字节的字符串,存储数组dtype名称(如'float32')
    fmt = '<ffffii10s'
    
    # 打包配置参数
    dtype_str = data.dtype.name.encode('utf-8').ljust(10, b'\x00')  # 补空字节至10位
    config_packed = struct.pack(
        fmt,
        x_range[0], x_range[1],
        y_range[0], y_range[1],
        xy_step,
        data.shape[0], data.shape[1],
        dtype_str
    )
    
    # 写入文件:先写struct配置,再写数组二进制数据
    with open(file_path, 'wb') as f:
        f.write(config_packed)
        data.tofile(f)
    
    # 可选:生成兼容pickle的索引文件(供原有读取程序过渡使用)
    pickle.dump(
        {'config_offset': 0, 'config_size': struct.calcsize(fmt), 'data_offset': struct.calcsize(fmt)},
        open(file_path + '.pkl', 'wb'),
        protocol=pickle.HIGHEST_PROTOCOL
    )

读取代码(兼容新旧格式)

import struct
import numpy as np
import pickle

def load_data(file_path):
    try:
        # 尝试用struct读取新格式
        fmt = '<ffffii10s'
        config_size = struct.calcsize(fmt)
        
        with open(file_path, 'rb') as f:
            # 读取配置块
            config_packed = f.read(config_size)
            x_start, x_end, y_start, y_end, xy_step, rows, cols, dtype_str = struct.unpack(fmt, config_packed)
            dtype = np.dtype(dtype_str.strip(b'\x00').decode('utf-8'))
            
            # 读取数组
            data = np.fromfile(f, dtype=dtype).reshape(rows, cols)
        
        return {
            'xyStep': xy_step,
            'xRange': (x_start, x_end),
            'yRange': (y_start, y_end),
            'data': data
        }
    except:
        # 兼容原有pickle格式
        with open(file_path, 'rb') as f:
            return pickle.load(f)

额外优化:使用np.memmap处理超大数组

如果扫描数据是逐步生成的(无需一次性加载到内存),可以用np.memmap直接写入磁盘,进一步降低内存占用:

def save_data_memmap(file_path, xy_step, x_range, y_range, shape, dtype=np.float32):
    # 先写入struct配置
    fmt = '<ffffii10s'
    dtype_str = dtype.name.encode('utf-8').ljust(10, b'\x00')
    config_packed = struct.pack(
        fmt,
        x_range[0], x_range[1],
        y_range[0], y_range[1],
        xy_step,
        shape[0], shape[1],
        dtype_str
    )
    
    with open(file_path, 'wb') as f:
        f.write(config_packed)
        # 预分配数组空间(写入空数据占位)
        f.seek(struct.calcsize(fmt) + np.prod(shape) * dtype.itemsize - 1)
        f.write(b'\x00')
    
    # 创建memmap对象,逐步写入数据
    memmap_obj = np.memmap(
        file_path,
        dtype=dtype,
        mode='r+',
        shape=shape,
        offset=struct.calcsize(fmt)
    )
    
    # 示例:模拟逐步写入扫描数据
    for i in range(shape[0]):
        memmap_obj[i] = np.random.rand(shape[1]).astype(dtype)  # 替换为实际扫描数据
    
    del memmap_obj  # 释放映射

关键注意事项

  • 字节序一致性:统一使用小端(<)格式,与numpy、pickle的默认存储字节序对齐,避免跨平台读取错误。
  • dtype固定:确保数组 dtype 与struct中存储的类型一致,读取时严格解析。
  • 兼容过渡:若原有读取程序无法修改,可生成配套的pickle索引文件,让原有程序通过索引读取.bin文件中的配置和数组。

内容的提问来源于stack exchange,提问作者Malum Phobos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.11 21:44:50