You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

读取大量图像数据时出现异常周期性写入的原因排查

问题描述

我在本地高速SSD中存储了约220,000张图像文件,使用Python及tifffile库将这些图像读取为numpy数组,合并为单个数组后保存至磁盘。读取该合并数组的速度远快于单独读取各文件,但在读取数据的过程中,出现了高达30MB/s的写入操作(预期流程应为:完成所有读取→创建合并数组→执行一次写入),且此时内存充足(整个数据集可完全放入内存)。我推测内存中存在未标记为已使用的数据,这解释了最初加载约15GB数据时未产生磁盘读取的现象。

相关代码

import os
import numpy as np
from tifffile import imread
from functools import partial
from tqdm.contrib.concurrent import process_map

def get_image(dir, ID):
    # 加载为numpy数组
    return imread(os.path.join(dir, ID + ".tif"))
    
def generate_numpy_file(IDs, folder, fname="train"):
    _read = partial(get_image, folder)

    print("Reading Data")
    images = process_map(_read, IDs, max_workers=20, chunksize=1024)
    images = np.array(images)

    print("Writing Data")
    np.save(os.path.join(SCRIPT_DIR, "Datasets", fname), images)
原因分析与解决方案

核心原因:多进程的数据传递机制

你使用的process_map基于多进程实现并行读取,而Python多进程间的数据传递依赖序列化(pickle)与管道通信:

  • 子进程读取的图像数组会被序列化,传递给主进程时,操作系统可能会将这些临时序列化数据写入磁盘swap分区或临时文件,这就是你观察到的意外写入操作。
  • 内存中未标记为已使用的数据,是多进程间数据拷贝留下的临时内存页,这也导致初期加载数据时系统未触发磁盘读取(临时数据仍在缓存中)。

解决方法

  • 改用线程池处理IO密集型任务
    图像读取属于IO密集型操作,线程池无需跨进程序列化数据,共享内存空间能避免额外磁盘写入:

    # 替换导入和调用
    from tqdm.contrib.concurrent import thread_map
    
    def generate_numpy_file(IDs, folder, fname="train"):
        _read = partial(get_image, folder)
        print("Reading Data")
        # 用thread_map替代process_map
        images = thread_map(_read, IDs, max_workers=20, chunksize=1024)
        images = np.array(images)
        print("Writing Data")
        np.save(os.path.join(SCRIPT_DIR, "Datasets", fname), images)
    
  • 预分配内存映射文件直接写入
    若坚持用多进程,可通过内存映射(memmap)让子进程直接写入目标文件,减少数据拷贝环节:

    import os
    import numpy as np
    from tifffile import imread
    from functools import partial
    from tqdm.contrib.concurrent import process_map
    
    def write_to_memmap(args):
        idx, folder, ID, memmap_path, shape, dtype = args
        img = imread(os.path.join(folder, ID + ".tif"))
        # 打开内存映射文件并写入对应位置
        memmap_arr = np.memmap(memmap_path, dtype=dtype, mode='r+', shape=shape)
        memmap_arr[idx] = img
        memmap_arr.flush()
        del memmap_arr
    
    def generate_numpy_file(IDs, folder, fname="train"):
        # 获取样本图像的形状和数据类型
        sample_img = imread(os.path.join(folder, IDs[0] + ".tif"))
        total_shape = (len(IDs),) + sample_img.shape
        dtype = sample_img.dtype
    
        save_path = os.path.join(SCRIPT_DIR, "Datasets", fname + ".npy")
        # 预创建内存映射文件
        memmap_arr = np.memmap(save_path, dtype=dtype, mode='w+', shape=total_shape)
        memmap_arr.flush()
        del memmap_arr
    
        print("Writing Data Directly")
        # 构造参数列表,传递给子进程
        args_list = [(idx, folder, ID, save_path, total_shape, dtype) for idx, ID in enumerate(IDs)]
        process_map(write_to_memmap, args_list, max_workers=20, chunksize=1024)
    
  • 临时关闭Swap(仅限内存充足场景)
    若不想修改代码,可临时关闭系统swap,阻止临时数据写入磁盘:

    # Linux系统执行
    sudo swapoff -a
    # 任务完成后恢复swap
    sudo swapon -a
    

内容的提问来源于stack exchange,提问作者Mandias

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 19:55:18