DeepFace人脸识别中find函数并行处理人脸的方案问询
基于DeepFace的多人脸识别并行优化方案
问题根源分析
ProcessPoolExecutor或原生多进程方案出现内存/显存溢出,核心原因是:
- 每个子进程会复制主进程中已加载的DeepFace模型,导致显存被重复占用;
- 未限制进程数,超出GPU显存承载上限;
- 多进程共享同一.pkl文件时可能触发IO锁,同时重复加载特征库会占用大量内存。
ThreadPoolExecutor能运行但性能不佳,是因为它属于CPU线程池,GPU推理任务会被串行调度,无法充分利用GPU并行计算能力。
优化方案
1. 优先采用批量处理(GPU效率最高)
GPU天生适合批量计算,相比多进程,将单帧中提取的所有人脸打包批量处理,能大幅降低显存开销与推理耗时。
实现代码示例:
# Recognition.py from deepface import DeepFace import pickle # 预加载特征库(仅加载一次) with open("face_db.pkl", "rb") as f: face_db = pickle.load(f) model = DeepFace.build_model("Facenet") def batch_recognize_faces(face_imgs): # 批量计算人脸特征 batch_embeddings = [DeepFace.represent(img_path=face, model=model) for face in face_imgs] # 批量比对特征库 results = [] for embedding in batch_embeddings: match = DeepFace.find(embedding=embedding, db_embeddings=face_db, enforce_detection=False) results.append(match) return results
# Main.py import cv2 from Recognition import batch_recognize_faces if __name__ == "__main__": frame = cv2.imread("test_frame.jpg") # 提取单帧中所有人脸 faces = [face["face"] for face in DeepFace.extract_faces(frame, enforce_detection=False)] # 批量识别 results = batch_recognize_faces(faces) print(results)
2. 多进程方案优化(需严格控制显存)
若必须用多进程,需确保每个子进程独立加载模型与特征库,同时限制进程数避免显存溢出。
实现代码示例:
# Recognition.py from deepface import DeepFace import pickle import torch # 子进程初始化函数:独立加载模型与特征库 def init_worker(pkl_path): global model, face_db # 加载模型(子进程独立实例) model = DeepFace.build_model("Facenet") # 加载专属.pkl副本(避免IO冲突) with open(pkl_path, "rb") as f: face_db = pickle.load(f) # 强制模型加载到GPU if torch.cuda.is_available(): model = model.cuda() def recognize_single_face(face_img): result = DeepFace.find( img_path=face_img, db_embeddings=face_db, model=model, enforce_detection=False ) # 清理当前进程显存碎片 torch.cuda.empty_cache() return result
# Main.py from concurrent.futures import ProcessPoolExecutor import cv2 from Recognition import init_worker, recognize_single_face import multiprocessing if __name__ == "__main__": # 预生成多份特征库副本(每个进程一份) pkl_paths = ["face_db_1.pkl", "face_db_2.pkl", "face_db_3.pkl"] # 限制进程数:根据GPU显存调整,例如16G显存最多开3-4个进程 max_workers = min(multiprocessing.cpu_count(), 3) # 初始化进程池,每个进程加载专属模型与特征库 with ProcessPoolExecutor( max_workers=max_workers, initializer=init_worker, initargs=(pkl_paths[0],) # 若有多个副本,可循环分配,此处简化为单副本示例 ) as executor: frame = cv2.imread("test_frame.jpg") faces = [face["face"] for face in DeepFace.extract_faces(frame, enforce_detection=False)] # 提交识别任务 futures = [executor.submit(recognize_single_face, face) for face in faces] results = [f.result() for f in futures] print(results)
3. 额外显存/内存优化技巧
- 使用轻量模型:替换
Facenet为OpenFace或DeepID,可降低单模型显存占用30%-50%; - 共享特征库:用
multiprocessing.Manager将特征库加载到共享内存,避免每个进程重复加载,大幅减少内存开销; - 显存清理:每次推理后调用
torch.cuda.empty_cache(),释放未使用的显存碎片; - 模型量化:对DeepFace模型进行INT8量化,进一步降低显存占用。
内容的提问来源于stack exchange,提问作者parsa
相关产品推荐
相关产品推荐

