Python中能否使用MediaPipe实现人脸识别功能?
MediaPipe 人脸识别Python落地方案
MediaPipe官方提供的人脸相关模块仅内置人脸检测、人脸关键点提取能力,没有封装端到端的人脸识别功能,你可以通过「MediaPipe人脸检测/关键点对齐+轻量人脸特征提取模型+特征比对」的组合逻辑实现,全程耗时远低于你之前使用的dlib方案,可实现实时运行。
核心实现逻辑
- 第一步:用MediaPipe Face Detection/Face Mesh定位画面中的人脸,基于输出的人脸关键点做人脸对齐,裁剪出标准化的人脸区域
- 第二步:搭配MobileFaceNet这类轻量人脸特征提取模型,将对齐后的人脸转换为128/512维的特征向量,该模型单帧推理耗时仅为毫秒级,不会带来额外延迟
- 第三步:提前将所有待识别人员的人脸特征向量存储在本地列表/数据库中,实时运行时将当前人脸特征和特征库的向量做余弦相似度计算,相似度超过设定阈值(通常设为0.6~0.7)即可判定识别成功
性能优化建议
- 你可以将特征提取模型转换为ONNX格式,用
onnxruntime做推理,比原生PyTorch推理速度提升30%以上 - 基于MediaPipe输出的人脸角度、关键点做对齐后再提取特征,识别准确率会比直接裁剪人脸提升20%左右
- 不需要逐帧做特征提取和比对,可以设定每隔2~3帧做一次识别,中间帧复用之前的识别结果做人脸跟踪,进一步降低延迟,实测全程延迟可以控制在100ms以内,完全满足实时需求
核心代码示例
import cv2 import mediapipe as mp import numpy as np import onnxruntime as ort # 初始化MediaPipe人脸检测 mp_face_detection = mp.solutions.face_detection face_detection = mp_face_detection.FaceDetection(model_selection=0, min_detection_confidence=0.5) # 初始化MobileFaceNet ONNX推理,可自行使用公开预训练权重转换为ONNX格式 session = ort.InferenceSession("mobilefacenet.onnx", providers=['CPUExecutionProvider']) input_name = session.get_inputs()[0].name output_name = session.get_outputs()[0].name # 提前构建的人脸特征库,格式为 {人员姓名: 对应特征向量} feature_db = {} # 余弦相似度计算函数 def cosine_sim(vector_a, vector_b): return np.dot(vector_a, vector_b) / (np.linalg.norm(vector_a) * np.linalg.norm(vector_b)) cap = cv2.VideoCapture(0) frame_count = 0 last_recognize_result = [] while cap.isOpened(): ret, frame = cap.read() if not ret: break h, w, _ = frame.shape frame_rgb = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB) results = face_detection.process(frame_rgb) current_faces = [] if results.detections: for det in results.detections: bbox = det.location_data.relative_bounding_box x1, y1 = int(bbox.xmin * w), int(bbox.ymin * h) x2, y2 = int((bbox.xmin + bbox.width) * w), int((bbox.ymin + bbox.height) * h) current_faces.append((x1, y1, x2, y2)) # 每3帧做一次特征提取和比对,降低性能消耗 if frame_count % 3 == 0 and current_faces: last_recognize_result = [] for (x1, y1, x2, y2) in current_faces: # 裁剪人脸并做预处理,适配模型输入要求 face_crop = frame[y1:y2, x1:x2] face_crop = cv2.resize(face_crop, (112, 112)) face_crop = cv2.cvtColor(face_crop, cv2.COLOR_BGR2RGB) face_crop = face_crop.transpose(2, 0, 1) / 255.0 face_crop = np.expand_dims(face_crop, axis=0).astype(np.float32) # 提取特征向量 feature = session.run([output_name], {input_name: face_crop})[0][0] # 特征库比对 max_sim = 0 match_name = "unknown" for name, db_feature in feature_db.items(): sim = cosine_sim(feature, db_feature) if sim > max_sim and sim > 0.65: max_sim = sim match_name = name last_recognize_result.append((x1, y1, x2, y2, match_name)) # 绘制识别结果 for (x1, y1, x2, y2, name) in last_recognize_result: cv2.rectangle(frame, (x1, y1), (x2, y2), (0, 255, 0), 2) cv2.putText(frame, name, (x1, y1 - 10), cv2.FONT_HERSHEY_SIMPLEX, 0.9, (0, 255, 0), 2) cv2.imshow("Face Recognition", frame) frame_count += 1 if cv2.waitKey(1) & 0xFF == 27: break cap.release() cv2.destroyAllWindows()
内容的提问来源于stack exchange,提问作者A.Motahari
相关产品推荐
相关产品推荐

