Python优化256MB二进制int32文件4字节序列读取速度的方法
高效读取二进制int32文件并排序的优化方案
你的两种方法慢的核心原因是Python层面的循环开销——不管是逐个读4字节还是手动切片,都绕不开Python循环的低效。下面是几种基于底层C实现的高效方案:
方案1:用array模块(轻量高效)
array模块专门针对同类型二进制数据设计,读取和写入都是底层优化的:
import array # 读取文件:'i'对应int32,若文件是特定字节序,可改用'>i'(大端)或'<i'(小端) with open('input.bin', 'rb') as f: num_arr = array.array('i') # 一次性读取指定数量的int32 num_arr.fromfile(f, numsPerFile) # 内置排序(同样是底层优化) num_arr.sort() # 写入排序后的文件 with open('output.bin', 'wb') as f: num_arr.tofile(f)
方案2:用struct模块批量解析
struct可以一次性把整个二进制缓冲区解析成int数组,避免Python循环:
import struct with open('input.bin', 'rb') as f: # 先读整个文件到缓冲区 buffer = f.read() # 批量解析所有int32,格式字符串里的'i'数量对应要读取的数字个数 num_tuple = struct.unpack(f'{numsPerFile}i', buffer) # 排序(转成列表后排序) sorted_nums = sorted(num_tuple) # 批量打包写入 with open('output.bin', 'wb') as f: f.write(struct.pack(f'{numsPerFile}i', *sorted_nums))
注:如果文件字节序和系统默认不同,要在格式字符串前加字节序标识,比如<{numsPerFile}i表示小端字节序。
方案3:用numpy(超大文件首选)
如果处理的文件更大,numpy的IO和排序效率会更突出:
import numpy as np # 读取文件,dtype指定int32,字节序可通过dtype设置(如'<i4') num_arr = np.fromfile('input.bin', dtype=np.int32, count=numsPerFile) # numpy内置排序,速度远快于Python原生排序 num_arr.sort() # 写入文件 num_arr.tofile('output.bin')
这几种方案的读取速度至少是你原来方法的10倍以上,因为完全避开了Python循环的开销,全部用底层C代码完成解析和操作。
内容的提问来源于stack exchange,提问作者Swif
相关产品推荐
相关产品推荐

