15MB二进制文件处理耗时过长,求Python代码优化方案
优化15MB二进制文件处理速度的方案
核心性能瓶颈
原代码耗时过长的主要原因:
- 频繁执行
del data[:frame]:删除列表开头元素会触发所有后续元素的内存移动,时间复杂度为O(n),累积开销极大。 - 纯Python循环逐字节处理:没有利用底层优化的批量计算能力,效率低下。
优化方案
方案1:避免修改原列表,用索引定位帧
直接通过索引访问原bytes对象,无需转列表或删除元素,切片访问效率远高于列表修改。
with open(r'C:\binary_file.bin', 'rb') as p: my_bytes = p.read() frame = 50 processed = [] scales = [ 0.0625, 0.0625, 0.00390625, 0.00390625, 3.05176e-05, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 0.0078125, 0.001953125, 0.0001220703, 0.0001220703, 0.0001220703, 3.05176e-05, 1.0, 0.0001220703, 0.001953125, 1.0, 1.0, 1.0 ] total_frames = len(my_bytes) // frame for sample in range(total_frames): start = sample * frame frame_data = my_bytes[start:start+frame] temp1 = [frame_data[i] + frame_data[i+1]*256 for i in range(0, frame-10, 2)] temp2 = [i - 65536 if i > 32767 else i for i in temp1] temp3 = [int(a*b) if b in (1.0, 2.0) else a*b for a, b in zip(temp2, scales)] temp3 = [round(i, 5) for i in temp3] temp3.insert(0, sample) processed.append(temp3)
关键改进:用切片代替列表删除,消除了最耗时的内存移动操作。
方案2:用struct批量解析16位整数
原代码手动计算16位整数的逻辑,可替换为C实现的struct模块批量解析,速度提升显著。
import struct with open(r'C:\binary_file.bin', 'rb') as p: my_bytes = p.read() frame = 50 # 每帧解析20个小端有符号16位整数,对应原逻辑的补码转换 fmt = f'<{int((frame-10)/2)}h' total_frames = len(my_bytes) // frame processed = [] scales = [ 0.0625, 0.0625, 0.00390625, 0.00390625, 3.05176e-05, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 0.0078125, 0.001953125, 0.0001220703, 0.0001220703, 0.0001220703, 3.05176e-05, 1.0, 0.0001220703, 0.001953125, 1.0, 1.0, 1.0 ] for sample in range(total_frames): start = sample * frame # 批量解析前40字节为16位整数 temp2 = struct.unpack(fmt, my_bytes[start:start+frame-10]) temp3 = [] for a, b in zip(temp2, scales): val = a * b temp3.append(int(val) if b in (1.0, 2.0) else round(val, 5)) temp3.insert(0, sample) processed.append(temp3)
关键改进:用struct替代手动整数转换,底层C实现比Python循环快10~20倍。
方案3:用numpy矢量化处理(最快方案)
利用numpy的矢量化数组运算,完全消除Python循环,处理15MB数据仅需几秒。
import numpy as np with open(r'C:\binary_file.bin', 'rb') as p: # 直接读取为uint8数组 data = np.frombuffer(p.read(), dtype=np.uint8) frame = 50 total_frames = data.size // frame # 重塑为(总帧数, 50)的二维数组 frames = data[:total_frames*frame].reshape(-1, frame) # 批量合并为16位有符号整数 temp2 = (frames[:, :40:2] + 256 * frames[:, 1:40:2]).astype(np.int16) # 应用缩放因子 scales_np = np.array(scales, dtype=np.float64) temp3 = temp2 * scales_np # 对缩放因子为1/2的情况转整数,其余保留5位小数 mask = (scales_np == 1.0) | (scales_np == 2.0) temp3[:, mask] = temp3[:, mask].astype(np.int64) temp3 = np.round(temp3, 5) # 添加sample列并转为列表 sample_col = np.arange(total_frames).reshape(-1, 1) processed = np.hstack([sample_col, temp3]).tolist()
关键改进:矢量化运算将多层Python循环替换为单次底层数组操作,性能提升最明显。
性能对比
- 原代码:耗时极长(数分钟级别)
- 方案1:耗时约为原代码的1/10~1/5
- 方案2:耗时约为原代码的1/20~1/10
- 方案3:耗时仅需几秒
内容的提问来源于stack exchange,提问作者sethdhanson
相关产品推荐
相关产品推荐

