You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

OpenCV中利用vector<Point>索引矩阵非零元素实现乘法优化

Optimizing Matrix Multiplication for Non-Zero Elements in OpenCV

Hey there! Let's fix that performance bottleneck you're dealing with. The core issue here is that your current full-matrix multiplication wastes cycles on zero-valued pixels, which we can avoid by using the vector<Point> of non-zero coordinates you already extracted. Let's break down the solutions step by step.

First: Fix the Incorrect Code in Your Example

Your attempt to use result.at<int>(nonZero) doesn't work because:

  • at<T>() expects a single Point (or row/col indices), not an entire vector.
  • Your matrix is CV_64FC1 (64-bit float), so you need to use double as the template type, not int.

Solution 1: Directly Iterate Over Non-Zero Coordinates

The simplest approach is to loop through each non-zero point and update only those pixels. This skips all zero-valued areas entirely, cutting down on unnecessary computations.

Option 1a: Start with a Clone of the Original Image

Since we only need to modify non-zero pixels, cloning the original image first lets us retain the zero values without extra work:

Mat img_temp(480, 640, CV_64FC1); 
Mat img = img_temp.clone(); 
Mat mask = Mat::ones(img.size(), CV_8UC1); 
double value = 3.56; 

// Apply mask
img_temp.copyTo(img, mask); 

// Tip: You can use mask directly here instead of img, since img's non-zeros match mask's
vector<Point> nonZero; 
findNonZero(mask, nonZero); 

// Efficient multiplication only on non-zero pixels
Mat result = img.clone();
for (const auto& pt : nonZero) {
    result.at<double>(pt) *= value;
}

Option 1b: Build from a Zero Matrix

If you need to start with an all-zero matrix (instead of cloning img), you can do this:

Mat result = Mat::zeros(img.size(), CV_64FC1);
for (const auto& pt : nonZero) {
    result.at<double>(pt) = img.at<double>(pt) * value;
}

Solution 2: Faster Memory Access with Pointers

For even better performance (especially with large datasets), skip the boundary checks of at<T>() and use direct memory pointers. This is ideal when you're sure your coordinates are valid (which they are, since they come from findNonZero):

Mat result = img.clone();
double* result_data = result.ptr<double>();
const double* img_data = img.ptr<double>();
int cols = img.cols;

for (const auto& pt : nonZero) {
    // Calculate the linear index of the pixel
    int linear_idx = pt.y * cols + pt.x;
    result_data[linear_idx] = img_data[linear_idx] * value;
}

Key Notes for Your Use Case

  • Reuse the Mask's Non-Zero Coordinates: Since your mask is the same across all 2000 matrices, you only need to run findNonZero(mask, nonZero) once, not for each matrix. This saves additional overhead.
  • Performance Gain: The smaller the percentage of non-zero pixels, the bigger the speedup compared to full-matrix multiplication. For sparse matrices, this can be orders of magnitude faster.

内容的提问来源于stack exchange,提问作者Mathieu Gauquelin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:49:23