如何用Python多进程处理大图像并将结果写入目标变量
超大图像Tile分片多进程优化方案
问题背景
处理25088×36864像素的超大图像,采用256×256像素的Tile分片方式处理,单进程耗时约61.36秒,且Windows任务管理器显示CPU、RAM、GPU或SSD利用率均未达50%,存在性能优化空间。需要实现多进程并行处理Tile,并安全将处理结果写入processedImage变量。
原单进程核心代码:
def processImage(self, img, tileSize = 256, numberOfThreads = 8): # a function within a class height, width, depth = img.shape print(height,width,depth,img.dtype) #create a duplicate but empty matrix same as the img processedImage = np.zeros((height,width,3), dtype=np.uint8) #calculate left and top offsets leftExcessPixels = int((width%tileSize)/2) topExcessPixels = int((height%tileSize)/2) #calculate the number of tiles columns(X) and row(Y) XNumberOfTiles = int(width/tileSize) YNumberOfTiles = int(height/tileSize) # for y in range(YNumberOfTiles): for x in range(XNumberOfTiles): XStart = (leftExcessPixels + (tileSize * x)) YStart = (topExcessPixels + (tileSize * y)) XEnd = XStart + tileSize YEnd = YStart + tileSize croppedImage = img[YStart:YEnd, XStart:XEnd] print('Y: ' + str(y) + ' X: ' + str(x),end=" ") #process the cropped images and store it on the same location on the empty image processedImage[YStart:YEnd, XStart:XEnd] = self.doSomeImageProcessing(croppedImage)
可行解决方案
方案1:进程池+结果集中写入
利用multiprocessing.Pool并行处理所有Tile,子进程返回处理后的Tile及对应坐标,主进程统一写入目标数组,完全避免多进程写入冲突,逻辑简单易维护。
修改后的完整代码:
from multiprocessing import Pool from time import monotonic import numpy as np class myClass(): def doSomeImageProcessing(self, args): # 解包任务参数 croppedImage, XStart, YStart, tileSize = args # 实际图像处理逻辑(示例为将Tile设为白色) processed_tile = croppedImage.copy() for i in range(255): processed_tile[:] = i+1 # 返回处理结果及坐标信息 return (XStart, YStart, tileSize, processed_tile) def processImage(self, tileSize = 256, num_processes=8): # 创建测试用超大图像 img = np.zeros((25088, 36864, 3), dtype=np.uint8) height, width, depth = img.shape print(height, width, depth, img.dtype) processedImage = np.zeros((height, width, depth), dtype=np.uint8) leftExcessPixels = int((width % tileSize)/2) topExcessPixels = int((height % tileSize)/2) XNumberOfTiles = int(width / tileSize) YNumberOfTiles = int(height / tileSize) # 生成所有Tile的任务参数列表 tasks = [] for y in range(YNumberOfTiles): for x in range(XNumberOfTiles): XStart = leftExcessPixels + tileSize * x YStart = topExcessPixels + tileSize * y croppedImage = img[YStart:YStart+tileSize, XStart:XStart+tileSize] tasks.append((croppedImage, XStart, YStart, tileSize)) print(f'Y: {y} X: {x}', end=" ") # 启动进程池并行处理 with Pool(num_processes) as pool: results = pool.map(self.doSomeImageProcessing, tasks) # 主进程统一写入处理结果 for XStart, YStart, tileSize, processed_tile in results: processedImage[YStart:YStart+tileSize, XStart:XStart+tileSize] = processed_tile # 验证处理结果 mean = np.mean(processedImage) if mean == 255: print(f'\nImage Processing successful: {mean}') else: print(f'\nImage Processing failed: {mean}') if __name__ == "__main__": x = myClass() start_time = monotonic() x.processImage(num_processes=8) print(f"Run time {monotonic() - start_time} seconds")
方案2:共享内存直接写入
针对超大规模图像,避免Tile数据拷贝带来的内存开销,使用共享内存让子进程直接写入目标区域。由于每个Tile的写入区域完全独立,无需加锁,性能更高。
修改后的完整代码:
from multiprocessing import Process, Array from time import monotonic import numpy as np class myClass(): def doSomeImageProcessing(self, args): # 解包共享内存及任务参数 shared_arr, img_shape, XStart, YStart, tileSize = args # 将共享内存转换为numpy数组 processedImage = np.frombuffer(shared_arr, dtype=np.uint8).reshape(img_shape) # 提取当前Tile并处理 cropped_tile = processedImage[YStart:YStart+tileSize, XStart:XStart+tileSize] for i in range(255): cropped_tile[:] = i+1 def processImage(self, tileSize = 256, num_processes=8): img = np.zeros((25088, 36864, 3), dtype=np.uint8) height, width, depth = img.shape print(height, width, depth, img.dtype) # 创建共享内存数组,存储处理后的图像('B'代表无符号字节) shared_arr = Array('B', height * width * depth, lock=False) # 转换为numpy数组并初始化 processedImage = np.frombuffer(shared_arr, dtype=np.uint8).reshape((height, width, depth)) processedImage[:] = 0 leftExcessPixels = int((width % tileSize)/2) topExcessPixels = int((height % tileSize)/2) XNumberOfTiles = int(width / tileSize) YNumberOfTiles = int(height / tileSize) # 创建并启动所有子进程 processes = [] for y in range(YNumberOfTiles): for x in range(XNumberOfTiles): XStart = leftExcessPixels + tileSize * x YStart = topExcessPixels + tileSize * y p = Process(target=self.doSomeImageProcessing, args=(shared_arr, (height, width, depth), XStart, YStart, tileSize)) processes.append(p) p.start() print(f'Y: {y} X: {x}', end=" ") # 等待所有子进程完成 for p in processes: p.join() # 验证处理结果 mean = np.mean(processedImage) if mean == 255: print(f'\nImage Processing successful: {mean}') else: print(f'\nImage Processing failed: {mean}') if __name__ == "__main__": x = myClass() start_time = monotonic() x.processImage(num_processes=8) print(f"Run time {monotonic() - start_time} seconds")
方案对比
- 方案1:逻辑简单,无同步问题,适合大多数场景;缺点是需要拷贝Tile数据,内存占用略高。
- 方案2:减少数据拷贝,内存效率更高,适合超大规模图像;需注意共享内存的使用细节,确保Tile区域无重叠。
内容的提问来源于stack exchange,提问作者Jake quin
相关产品推荐
相关产品推荐

