如何加速Python中加密图像文件的解包操作?
我从加密文件中提取2048x2048图像及相关信息,单张加载耗时约1.7秒。但每次迭代要加载约40张图,后续迭代次数还会增加,必须优化速度。试过PyPy、Numba这些JIT工具:
- PyPy解包只花0.3秒,但调用numpy的reshape操作时耗时直接翻倍;
- CPython里用Numba的
@njit或@vectorize装饰器时,因为unpack函数不被支持,报TypingError;
还没试Cython,想知道怎么让PyPy或Numba生效,或者最优的加速方案,也愿意尝试C++优化,但不知道怎么和Python衔接。
相关代码:
from struct import unpack import numpy as np def read_file(filename: str, nx: int, ny: int) -> tuple: f = open(filename, "rb") raw = [unpack('d', f.read(8))[0] for _ in range(2*nx*ny)] #Creates 1D list real_image = np.asarray(raw[0::2]).reshape(nx,ny) #Every other point is the real part of the image imaginary_image = np.asarray(raw[1::2]).reshape(nx,ny) #Every other point +1 is imaginary part of image return real_image, imaginary_image
Numba报错信息:
File c:\Users\MyName\Anaconda3\envs\myenv\Lib\site-packages\numba\core\dispatcher.py:468, in _DispatcherBase._compile_for_args(self, *args, **kws) 464 msg = (f"{str(e).rstrip()} \n\nThis error may have been caused " 465 f"by the following argument(s):\n{args_str}\n") 466 e.patch_message(msg) --> 468 error_rewrite(e, 'typing') 469 except errors.UnsupportedError as e: 470 # Something unsupported is present in the user code, add help info 471 error_rewrite(e, 'unsupported_error') File c:\Users\MyName\Anaconda3\envs\myenv\Lib\site-packages\numba\core\dispatcher.py:409, in _DispatcherBase._compile_for_args.<locals>.error_rewrite(e, issue_type) 407 raise e 408 else: --> 409 raise e.with_traceback(None) TypingError: Failed in nopython mode pipeline (step: nopython frontend) Untyped global name 'unpack': Cannot determine Numba type of <class 'builtin_function_or_method'>
一、原生Python优化(最快见效)
你的代码最大性能瓶颈是逐次调用unpack和f.read(8),循环800多万次完全没必要。直接一次性读取所有字节,用numpy批量解析,能把单张加载时间压到毫秒级:
import numpy as np def read_file_optimized(filename: str, nx: int, ny: int) -> tuple: # 计算总字节数:每个double占8字节,共2*nx*ny个元素 total_bytes = 2 * nx * ny * 8 with open(filename, "rb") as f: raw_data = f.read(total_bytes) # 批量解析字节为numpy数组,dtype=np.float64对应struct的'd'格式 arr = np.frombuffer(raw_data, dtype=np.float64) # 直接拆分实部虚部并reshape,全程numpy C级操作 real_image = arr[0::2].reshape(nx, ny) imaginary_image = arr[1::2].reshape(nx, ny) return real_image, imaginary_image
这个版本完全避开Python循环,性能比原代码提升几十倍,单张加载时间能降到0.1秒以内,CPython和PyPy下都能高效运行。
二、让Numba生效的方案
Numba不支持struct.unpack,但可以通过两种方式适配:
方案1:IO操作放外部,Numba处理数组
把文件读取逻辑留在Python层,只让Numba加速数组拆分和reshape:
import numpy as np from numba import njit def read_file_numba(filename: str, nx: int, ny: int) -> tuple: total_bytes = 2 * nx * ny * 8 with open(filename, "rb") as f: raw_data = f.read(total_bytes) arr = np.frombuffer(raw_data, dtype=np.float64) return split_and_reshape(arr, nx, ny) @njit def split_and_reshape(arr, nx, ny): real = arr[0::2].reshape(nx, ny) imag = arr[1::2].reshape(nx, ny) return real, imag
不过这个优化意义不大,因为numpy本身已经是C级实现,Numba在这里提升有限。
方案2:Numba直接用低级IO读取
Numba支持os.open/os.read这类系统级IO函数,可以把整个流程放到Numba编译函数里:
import numpy as np from numba import njit import os @njit def read_file_numba_direct(filename: str, nx: int, ny: int) -> tuple: fd = os.open(filename, os.O_RDONLY) total_bytes = 2 * nx * ny * 8 raw_data = os.read(fd, total_bytes) os.close(fd) arr = np.frombuffer(raw_data, dtype=np.float64) real = arr[0::2].reshape(nx, ny) imag = arr[1::2].reshape(nx, ny) return real, imag
适合需要更复杂字节处理的场景,性能和原生numpy优化版差距不大。
三、PyPy优化方案
PyPy对numpy reshape性能差,是因为其numpy兼容层(cpyext)对部分操作支持不足。直接用上面的原生Python优化版即可——该版本的numpy操作都是最基础的frombuffer、切片和reshape,PyPy对这些操作的支持已经很完善,性能和CPython接近甚至更快。
另外,PyPy下绝对要避免循环调用Python内置函数(比如原代码的unpack循环),改用批量操作就能解决核心性能问题。
四、C++衔接方案(极致性能需求)
如果以上方案还不够,可以用C++写读取逻辑,通过两种方式和Python衔接:
1. 使用ctypes
把C代码编译成动态链接库,用Python的ctypes调用:
C代码(read_image.cpp):
#include <fstream> #include <vector> #include <cstdint> extern "C" { void read_file(const char* filename, double* real_out, double* imag_out, int nx, int ny) { std::ifstream file(filename, std::ios::binary); int total_elements = 2 * nx * ny; std::vector<double> data(total_elements); file.read(reinterpret_cast<char*>(data.data()), total_elements * sizeof(double)); for (int i = 0; i < nx * ny; ++i) { real_out[i] = data[2*i]; imag_out[i] = data[2*i + 1]; } } }
Windows编译命令:g++ -shared -o read_image.dll read_image.cpp
Python调用代码:
import ctypes import numpy as np def read_file_cpp(filename: str, nx: int, ny: int) -> tuple: lib = ctypes.CDLL("./read_image.dll") read_func = lib.read_file read_func.argtypes = [ctypes.c_char_p, ctypes.POINTER(ctypes.c_double), ctypes.POINTER(ctypes.c_double), ctypes.c_int, ctypes.c_int] real_image = np.zeros((nx, ny), dtype=np.float64) imaginary_image = np.zeros((nx, ny), dtype=np.float64) read_func(filename.encode('utf-8'), real_image.ctypes.data_as(ctypes.POINTER(ctypes.c_double)), imaginary_image.ctypes.data_as(ctypes.POINTER(ctypes.c_double)), nx, ny) return real_image, imaginary_image
2. 使用pybind11
pybind11是更便捷的C++/Python绑定工具,无需手动处理类型转换:
C++代码(read_image.cpp):
#include <pybind11/pybind11.h> #include <pybind11/numpy.h> #include <fstream> namespace py = pybind11; py::tuple read_file(const std::string& filename, int nx, int ny) { std::ifstream file(filename, std::ios::binary); int total_elements = 2 * nx * ny; py::array_t<double> data(total_elements); file.read(reinterpret_cast<char*>(data.mutable_data()), total_elements * sizeof(double)); auto data_ptr = data.mutable_data(); py::array_t<double> real_image({nx, ny}); py::array_t<double> imag_image({nx, ny}); auto real_ptr = real_image.mutable_data(); auto imag_ptr = imag_image.mutable_data(); for (int i = 0; i < nx * ny; ++i) { real_ptr[i] = data_ptr[2*i]; imag_ptr[i] = data_ptr[2*i + 1]; } return py::make_tuple(real_image, imag_image); } PYBIND11_MODULE(read_image, m) { m.def("read_file", &read_file, "Read real and imaginary images from binary file"); }
用setup.py编译:
from setuptools import setup, Extension import pybind11 ext_modules = [ Extension( "read_image", ["read_image.cpp"], include_dirs=[pybind11.get_include()], language='c++' ), ] setup( name="read_image", ext_modules=ext_modules, setup_requires=['pybind11>=2.6.0'], )
编译后直接在Python中导入使用:import read_image; read_image.read_file(filename, nx, ny)
内容的提问来源于stack exchange,提问作者Alex Tripsas

