C++下Dlib人脸检测性能远逊于Python的优化方案咨询
C++下Dlib人脸检测性能不及Python的优化方案
问题描述
我在C中使用Dlib进行人脸检测时性能远不如Python:单帧检测在C中耗时约1-2秒,而Python中仅需500ms以内。已按说明在编译库和代码时启用DUSE_AVX_INSTRUCTIONS,并添加编译参数-mavx2,虽比之前的3-4秒有所改善,但仍未达预期。
迁移到C++是为了让程序在树莓派4B这类迷你电脑上更快速、轻量化,但目前性能未达预期,求解决方案。
开发环境
- OS:Ubuntu 22.04
- 处理器:Intel(R) Core(TM) i5-7200U CPU @ 2.50GHz
- 内存:8GB
- GPU:NVIDIA 930MX(安装OpenCV和Dlib时未启用)
现有C++代码
#include "opencv2/objdetect.hpp" #include "opencv2/highgui.hpp" #include "opencv2/imgproc.hpp" #include <dlib/opencv.h> #include <dlib/image_processing/frontal_face_detector.h> #include <dlib/image_processing.h> #include <chrono> using namespace dlib; using namespace std; int main() { cv::VideoCapture cap(0); std::vector<cv::Rect> facesCV; std::vector<rectangle> faces; frontal_face_detector detector = get_frontal_face_detector(); cv::namedWindow("test"); cv::Mat frame, small; if (!cap.isOpened()) { cerr << "Unable to connect to camera" << endl; return 1; } while (true) { // Grab a frame if (!cap.read(frame)) { break; } cv::resize(frame, small, {640, 480}); cv_image<rgb_pixel> cimg(small); auto start = std::chrono::high_resolution_clock::now(); // Detect faces faces = detector(cimg); for (auto &f : faces) { facesCV.emplace_back(cv::Point((int)f.left(), (int)f.top()), cv::Point((int)f.right(), (int)f.bottom())); } for (auto &r : facesCV) { cv::rectangle(small, r, {0, 255, 0}, 2); } auto end = std::chrono::high_resolution_clock::now(); std::chrono::duration<double, std::milli> elapsed = end - start; cout << "Time taken for processing: " << elapsed.count() << " ms" << endl; cv::imshow("test", small); cv::waitKey(1); faces.clear(); facesCV.clear(); } }
现有编译命令
sudo g++ -g test.cpp -o dlib_test `pkg-config --cflags --libs opencv4` -ldlib -lX11 -lblas -llapack -mavx2
优化解决方案
1. 调整编译参数,启用最高级优化
当前编译命令使用-g(调试模式)会关闭大量优化,需移除并添加-O3(最高级编译优化);同时i5-7200U支持FMA指令集,可添加-mfma进一步提升性能。修改后的编译命令:
g++ -O3 test.cpp -o dlib_test `pkg-config --cflags --libs opencv4` -ldlib -lX11 -lblas -llapack -mavx2 -mfma
注意:无需使用sudo编译普通程序,避免权限问题。
2. 确保Dlib库已正确启用AVX/FMA优化
重新编译Dlib时,需明确指定CMake参数确认优化生效:
cmake -DDLIB_USE_AVX_INSTRUCTIONS=ON -DDLIB_USE_FMA_INSTRUCTIONS=ON .. make -j$(nproc) sudo make install
编译过程中需看到控制台输出Enabling AVX instructions和Enabling FMA instructions,才说明优化已生效。
3. 简化图像处理与数据转换流程
- 跳过中间容器拷贝:直接用检测结果绘制矩形,移除
facesCV的创建与拷贝步骤:// 替换原有的两个for循环 for (auto &f : faces) { cv::rectangle(small, cv::Rect(f.left(), f.top(), f.width(), f.height()), {0, 255, 0}, 2); } - 直接设置相机分辨率,跳过resize操作:
cap.set(cv::CAP_PROP_FRAME_WIDTH, 640); cap.set(cv::CAP_PROP_FRAME_HEIGHT, 480); // 移除循环内的cv::resize(frame, small, {640, 480}); - 若对图像质量要求不高,resize时使用更快的插值方式:
cv::resize(frame, small, {640, 480}, 0, 0, cv::INTER_NEAREST);
4. 切换到更轻量的检测器
Dlib默认的HOG检测器精度高但速度一般,可尝试使用CNN检测器(需下载预训练模型mmod_human_face_detector.dat):
// 替换原有的detector初始化 dlib::cnn_face_detection_model_v1 detector("mmod_human_face_detector.dat"); // 检测逻辑不变 auto faces = detector(cimg);
CNN检测器在CPU上的性能表现优于HOG,尤其适合小尺寸图像。
5. 启用OpenCV硬件加速
针对树莓派4B,编译OpenCV时需添加ARM架构优化参数:-DWITH_V4L=ON和-DWITH_NEON=ON;代码中使用V4L2接口打开相机提升捕获效率:
cv::VideoCapture cap(0, cv::CAP_V4L2);
内容的提问来源于stack exchange,提问作者Muhammad Iqbal
相关产品推荐
相关产品推荐

