OpenCV中利用vector<Point>索引矩阵非零元素实现乘法优化
Hey there! Let's fix that performance bottleneck you're dealing with. The core issue here is that your current full-matrix multiplication wastes cycles on zero-valued pixels, which we can avoid by using the vector<Point> of non-zero coordinates you already extracted. Let's break down the solutions step by step.
First: Fix the Incorrect Code in Your Example
Your attempt to use result.at<int>(nonZero) doesn't work because:
at<T>()expects a singlePoint(or row/col indices), not an entire vector.- Your matrix is
CV_64FC1(64-bit float), so you need to usedoubleas the template type, notint.
Solution 1: Directly Iterate Over Non-Zero Coordinates
The simplest approach is to loop through each non-zero point and update only those pixels. This skips all zero-valued areas entirely, cutting down on unnecessary computations.
Option 1a: Start with a Clone of the Original Image
Since we only need to modify non-zero pixels, cloning the original image first lets us retain the zero values without extra work:
Mat img_temp(480, 640, CV_64FC1); Mat img = img_temp.clone(); Mat mask = Mat::ones(img.size(), CV_8UC1); double value = 3.56; // Apply mask img_temp.copyTo(img, mask); // Tip: You can use mask directly here instead of img, since img's non-zeros match mask's vector<Point> nonZero; findNonZero(mask, nonZero); // Efficient multiplication only on non-zero pixels Mat result = img.clone(); for (const auto& pt : nonZero) { result.at<double>(pt) *= value; }
Option 1b: Build from a Zero Matrix
If you need to start with an all-zero matrix (instead of cloning img), you can do this:
Mat result = Mat::zeros(img.size(), CV_64FC1); for (const auto& pt : nonZero) { result.at<double>(pt) = img.at<double>(pt) * value; }
Solution 2: Faster Memory Access with Pointers
For even better performance (especially with large datasets), skip the boundary checks of at<T>() and use direct memory pointers. This is ideal when you're sure your coordinates are valid (which they are, since they come from findNonZero):
Mat result = img.clone(); double* result_data = result.ptr<double>(); const double* img_data = img.ptr<double>(); int cols = img.cols; for (const auto& pt : nonZero) { // Calculate the linear index of the pixel int linear_idx = pt.y * cols + pt.x; result_data[linear_idx] = img_data[linear_idx] * value; }
Key Notes for Your Use Case
- Reuse the Mask's Non-Zero Coordinates: Since your mask is the same across all 2000 matrices, you only need to run
findNonZero(mask, nonZero)once, not for each matrix. This saves additional overhead. - Performance Gain: The smaller the percentage of non-zero pixels, the bigger the speedup compared to full-matrix multiplication. For sparse matrices, this can be orders of magnitude faster.
内容的提问来源于stack exchange,提问作者Mathieu Gauquelin

