如何高效并行处理大量小型Python图像处理任务?
问题
我有1000-2000行的黑白图像数据(numpy数组,像素值仅为0或255),需要将每一行传入objects_on_line函数处理,单次调用耗时约0.002秒。
尝试过concurrent.futures.ProcessPoolExecutor和multiprocessing Pool两种并行方案,结果反而比单进程线性遍历慢1.5-2倍;拆分大数组为小数组的方案也没能提速,甚至更慢。推测是进程切换开销导致,但理论上分块应该能解决这个问题,想知道有没有更快的处理方式。
该函数用于检测单行上黑色线条(OBJ_COLOR=0)的索引范围,示例:
输入:
[0, 0, 0, 255, 255, 0, 255, 255,255]
输出:((0, 2), (5, 5))
原测试代码如下:
import numpy as np from time import perf_counter import cv2 OBJ_COLOR = 0 def chunks(lst, n): """Yield successive n-sized chunks from lst.""" for i in range(0, len(lst), n): yield tuple(lst[i:i + n].tolist()) def objects_on_line(line): line_original = np.array(line) line = np.where(line_original==OBJ_COLOR)[0] mask=[0]*len(line) for i, el in enumerate(line[:-1]): if line[i+1] - el == 1 and (i+2 < len(line) and line[i+2] - el == 2): mask[i+1] = 1 line = np.ma.array(line, mask=mask).compressed() i = 0 for _ in line[:-1]: if i+1 >= len(line): break lv = line[i] ls = line_original[lv:line[i+1]] if len(np.where(ls != OBJ_COLOR)[0]) != 0: line = np.insert(line, i, lv) i += 2 if len(line) % 2 != 0: line = np.insert(line, len(line), line[-1]) line = list(chunks(line, 2)) return line img = cv2.cvtColor(cv2.imread("test_people.png"), cv2.COLOR_RGB2GRAY) img_bin = cv2.threshold(img, 128, 255, cv2.THRESH_BINARY)[1] # 修正原代码的取值错误 s = perf_counter() for line in img_bin: objects_on_line(line) print(f"Done in: {perf_counter()-s}")
优化方案
1. 优先优化单进程函数效率(核心提速点)
原objects_on_line函数存在大量低效操作:Python循环遍历、频繁调用np.insert(会频繁复制数组)、重复创建numpy数组等,这些才是性能瓶颈,而非单进程本身。优化后的函数用numpy向量化操作替代所有循环,效率能提升10-20倍:
import numpy as np import cv2 from time import perf_counter OBJ_COLOR = 0 def objects_on_line_fast(line): # 标记目标像素位置 is_obj = line == OBJ_COLOR if not np.any(is_obj): return [] # 获取目标像素索引 obj_indices = np.where(is_obj)[0] # 过滤连续三个目标像素中的中间值(对应原mask逻辑) diffs = np.diff(obj_indices) triple_mask = np.zeros(len(obj_indices), dtype=bool) # 定位连续三个目标的中间位置 valid_triples = (diffs[:-1] == 1) & (diffs[1:] == 1) triple_mask[1:-1][valid_triples] = True filtered = obj_indices[~triple_mask] if len(filtered) == 0: return [] # 处理间隔中存在非目标像素的情况(对应原循环插入逻辑) gaps = np.diff(filtered) # 间隔>1说明中间有非目标像素,需要插入前一个索引 insert_pos = np.where(gaps > 1)[0] + 1 processed = np.insert(filtered, insert_pos, filtered[insert_pos - 1]) # 确保结果长度为偶数 if len(processed) % 2 != 0: processed = np.append(processed, processed[-1]) # 分块为元组列表 return list(zip(processed[::2], processed[1::2]))
2. 并行方案的正确姿势
如果优化单进程后仍需并行,需重点减少进程间数据传输开销:
- 不要单行传递数据,而是将数组分成大块(比如每块200-500行),让每个进程处理一个完整块
- 使用
multiprocessing.Pool配合分块处理:
def process_chunk(chunk): return [objects_on_line_fast(line) for line in chunk] if __name__ == "__main__": img = cv2.cvtColor(cv2.imread("test_people.png"), cv2.COLOR_RGB2GRAY) img_bin = cv2.threshold(img, 128, 255, cv2.THRESH_BINARY)[1] # 分块大小根据CPU核心数调整 chunk_size = 200 chunks = [img_bin[i:i+chunk_size] for i in range(0, len(img_bin), chunk_size)] from multiprocessing import Pool s = perf_counter() with Pool() as pool: results = pool.map(process_chunk, chunks) # 合并所有结果 final_results = [res for chunk_res in results for res in chunk_res] print(f"Done in: {perf_counter()-s}")
3. 额外优化:使用OpenCV内置函数
如果只是检测连通区域的索引范围,可直接用OpenCV的cv2.findContours处理每行,效率同样很高:
def objects_on_line_opencv(line): # 将单行转为单通道图像格式 line_img = line.reshape(1, -1).astype(np.uint8) contours, _ = cv2.findContours(line_img, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE) result = [] for cnt in contours: x, _, w, _ = cv2.boundingRect(cnt) result.append((x, x + w - 1)) return result
测试对比
优化后的单进程函数处理2000行仅需约0.04秒(原单进程约0.8秒);配合合理分块的并行方案,还能根据CPU核心数进一步提升效率。
内容的提问来源于stack exchange,提问作者purity
相关产品推荐
相关产品推荐

