You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python优化256MB二进制int32文件4字节序列读取速度的方法

高效读取二进制int32文件并排序的优化方案

你的两种方法慢的核心原因是Python层面的循环开销——不管是逐个读4字节还是手动切片,都绕不开Python循环的低效。下面是几种基于底层C实现的高效方案:

方案1:用array模块(轻量高效)

array模块专门针对同类型二进制数据设计,读取和写入都是底层优化的:

import array

# 读取文件:'i'对应int32,若文件是特定字节序,可改用'>i'(大端)或'<i'(小端)
with open('input.bin', 'rb') as f:
    num_arr = array.array('i')
    # 一次性读取指定数量的int32
    num_arr.fromfile(f, numsPerFile)

# 内置排序(同样是底层优化)
num_arr.sort()

# 写入排序后的文件
with open('output.bin', 'wb') as f:
    num_arr.tofile(f)

方案2:用struct模块批量解析

struct可以一次性把整个二进制缓冲区解析成int数组,避免Python循环:

import struct

with open('input.bin', 'rb') as f:
    # 先读整个文件到缓冲区
    buffer = f.read()
    # 批量解析所有int32,格式字符串里的'i'数量对应要读取的数字个数
    num_tuple = struct.unpack(f'{numsPerFile}i', buffer)

# 排序(转成列表后排序)
sorted_nums = sorted(num_tuple)

# 批量打包写入
with open('output.bin', 'wb') as f:
    f.write(struct.pack(f'{numsPerFile}i', *sorted_nums))

注:如果文件字节序和系统默认不同,要在格式字符串前加字节序标识,比如<{numsPerFile}i表示小端字节序。

方案3:用numpy(超大文件首选)

如果处理的文件更大,numpy的IO和排序效率会更突出:

import numpy as np

# 读取文件,dtype指定int32,字节序可通过dtype设置(如'<i4')
num_arr = np.fromfile('input.bin', dtype=np.int32, count=numsPerFile)

# numpy内置排序,速度远快于Python原生排序
num_arr.sort()

# 写入文件
num_arr.tofile('output.bin')

这几种方案的读取速度至少是你原来方法的10倍以上,因为完全避开了Python循环的开销,全部用底层C代码完成解析和操作。

内容的提问来源于stack exchange,提问作者Swif

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 10:41:18