You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python多线程磁盘I/O优化:华硕Hyper M.2卡性能测试代码改进咨询

优化磁盘IO测试代码:从多线程随机到顺序执行提升性能

你的测试结果和CrystalDiskMark随机读写接近,说明当前多线程逻辑确实是在模拟随机IO场景。要提升性能并减少随机性,完全可以通过调整为顺序执行读写(或结合批量顺序策略)来实现,尤其是NVMe设备对顺序IO的带宽利用率远高于小文件随机IO。

核心优化思路

  • 减少文件系统元数据开销:多线程并发创建、读写大量小文件会频繁触发文件系统的目录项更新、inode分配,这会成为性能瓶颈,顺序执行则能降低这类竞争。
  • 消除非磁盘瓶颈:原代码用os.urandom生成测试数据,这个操作依赖系统熵池,速度有限且不稳定,会掩盖真实的磁盘性能。
  • 匹配NVMe特性:NVMe的顺序读写带宽远高于随机小文件读写,调整测试逻辑能更准确地测出设备的极限性能。

优化后的代码(保持多文件+顺序执行)

import os
import time
import logging
import datetime
import shutil

def write_file(file_path, size_mb):
    # 用固定字节流代替os.urandom,消除随机数生成瓶颈
    data = b'\x00' * (size_mb * 1024 * 1024)
    with open(file_path, 'wb') as f:
        f.write(data)

def read_file(file_path):
    with open(file_path, 'rb') as f:
        # 分块读取,避免一次性加载大文件到内存
        while f.read(1024*1024):
            pass

def simulate_disk_io(base_path, num_files, size_mb):
    os.makedirs(base_path, exist_ok=True)
    file_paths = [os.path.join(base_path, f"file_{i}.dat") for i in range(num_files)]
    
    # 顺序写入所有文件
    start_time = time.time()
    logging.info(f"TIME: {datetime.datetime.now()}, WRITING FILES STARTED (SEQUENTIAL)")
    for path in file_paths:
        write_file(path, size_mb)
    write_duration = time.time() - start_time
    logging.info(f"Writing {num_files} files took {write_duration:.2f} seconds.")
    logging.info(f"Write Throughput: {num_files * size_mb / write_duration:.2f} MB/s")
    
    # 顺序读取所有文件
    start_time = time.time()
    logging.info(f"TIME: {datetime.datetime.now()}, READING FILES STARTED (SEQUENTIAL)")
    for path in file_paths:
        read_file(path)
    read_duration = time.time() - start_time
    logging.info(f"Reading {num_files} files took {read_duration:.2f} seconds.")
    logging.info(f"Read Throughput: {num_files * size_mb / read_duration:.2f} MB/s")
    
    shutil.rmtree(base_path)
    logging.info(f"TIME: {datetime.datetime.now()}, Cleanup completed.")

if __name__ == "__main__":
    # 修复logging配置的参数错误
    logging.basicConfig(filename='read_write_test.log', level=logging.INFO, 
                        format='%(asctime)s - %(levelname)s - %(message)s')
    simulate_disk_io(r"F:\\disk_test", 5000, 11)

另一种方案:用单个大文件模拟总数据量(更贴近顺序IO极限)

如果不需要测试多文件场景,直接生成一个总大小为5000*11=55000MB的大文件,能更准确测出NVMe的顺序读写极限:

import os
import time
import logging
import datetime
import shutil

def write_large_file(file_path, total_size_mb):
    chunk_size = 1024 * 1024 * 64  # 64MB块,匹配磁盘IO块大小
    data = b'\x00' * chunk_size
    total_chunks = total_size_mb // 64
    remaining = total_size_mb % 64
    
    with open(file_path, 'wb') as f:
        for _ in range(total_chunks):
            f.write(data)
        if remaining > 0:
            f.write(b'\x00' * (remaining * 1024 * 1024))

def read_large_file(file_path):
    chunk_size = 1024 * 1024 * 64
    with open(file_path, 'rb') as f:
        while f.read(chunk_size):
            pass

def simulate_large_file_io(base_path, total_size_mb):
    os.makedirs(base_path, exist_ok=True)
    file_path = os.path.join(base_path, "large_test_file.dat")
    
    start_time = time.time()
    logging.info(f"TIME: {datetime.datetime.now()}, WRITING LARGE FILE STARTED")
    write_large_file(file_path, total_size_mb)
    write_duration = time.time() - start_time
    logging.info(f"Writing {total_size_mb}MB file took {write_duration:.2f} seconds.")
    logging.info(f"Write Throughput: {total_size_mb / write_duration:.2f} MB/s")
    
    start_time = time.time()
    logging.info(f"TIME: {datetime.datetime.now()}, READING LARGE FILE STARTED")
    read_large_file(file_path)
    read_duration = time.time() - start_time
    logging.info(f"Reading {total_size_mb}MB file took {read_duration:.2f} seconds.")
    logging.info(f"Read Throughput: {total_size_mb / read_duration:.2f} MB/s")
    
    shutil.rmtree(base_path)
    logging.info(f"TIME: {datetime.datetime.now()}, Cleanup completed.")

if __name__ == "__main__":
    logging.basicConfig(filename='large_file_test.log', level=logging.INFO, 
                        format='%(asctime)s - %(levelname)s - %(message)s')
    simulate_large_file_io(r"F:\\disk_test", 5000*11)

关键改动说明

  1. 替换随机数据生成:用b'\x00'填充数据,避免os.urandom的性能瓶颈,确保测试的是磁盘真实性能。
  2. 顺序执行读写:去掉多线程,改为单线程顺序处理文件,消除文件系统元数据竞争,提升IO效率。
  3. 分块读取文件:避免一次性将大文件加载到内存,减少内存占用,同时更贴近真实磁盘IO模式。
  4. 修复原代码bug:补全datetime导入,修正logging.basicConfig的参数格式,确保日志正常输出。

内容的提问来源于stack exchange,提问作者Gökce

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 14:19:56