You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pybind11绑定C++ cv::Mat返回Python端图像像素错位问题

问题描述

我需要编写一个接收图像作为入参、返回处理后图像的函数,使用pybind11将该函数绑定到Python环境调用。
目前Python端传图像到C的逻辑已经跑通,但C处理完的图像返回Python环节出现异常。
以下是实现图像翻转功能的示例代码:

py::array_t<uint8_t> flipcvMat(py::array_t<uint8_t>& img)
{
    auto rows = img.shape(0);
    auto cols = img.shape(1);
    auto channels = img.shape(2);
    std::cout << "rows: " << rows << " cols: " << cols << " channels: " << channels << std::endl;
    auto type = CV_8UC3;

    cv::Mat cvimg2(rows, cols, type, (unsigned char*)img.data());

    cv::imwrite("/source/test.png", cvimg2); // 本地写入正常

    cv::Mat cvimg3(rows, cols, type);
    cv::flip(cvimg2, cvimg3, 0);

    cv::imwrite("/source/testout.png", cvimg3); // 本地写入正常

    py::array_t<uint8_t> output(
                                py::buffer_info(
                                cvimg3.data,
                                sizeof(uint8_t), //itemsize
                                py::format_descriptor<uint8_t>::format(),
                                3, // ndim
                                std::vector<size_t> {rows, cols , 3}, // shape
                                std::vector<size_t> {cols * sizeof(uint8_t), sizeof(uint8_t), 3} // strides
    )
    );
    return output;
}

Python端调用代码:

import cv2
img = cv2.imread('/source/whatever/ubuntu-1.png')
img3= opencvtest.flipcvMat(img)

已知验证结果:

  • 输入为3通道图像
  • C++侧通过cv::imwrite保存的输入副本、翻转后结果图像均正常,翻转逻辑本身没有问题
    故障现象:Python端接收到的返回图像存在像素错位,只能模糊识别轮廓,像素排列完全混乱。初步判断是构造py::buffer_info时参数配置错误,但未定位到具体问题。
问题根因

核心错误是py::buffer_info中strides步长参数配置错误,另外还隐藏了内存生命周期的野指针风险。
步长的定义是:从当前维度的一个元素移动到同维度下一个相邻元素,需要跨过的字节总数。你当前写的步长{cols * sizeof(uint8_t), sizeof(uint8_t), 3}完全不符合HWC格式三通道图像的内存排布规则:

  • 第0维对应图像行(高):从一行首像素移动到下一行首像素,需要跨过整行所有像素的所有通道,步长应为cols * channels * sizeof(uint8_t),你写的步长没有乘通道数,相当于每读一行只跳了1/3的行长度,必然错位
  • 第1维对应图像列(宽):从某列像素移动到下一列同位置像素,需要跨过1个像素的全部通道,步长应为channels * sizeof(uint8_t),你写的步长只跳1字节,每次只挪到同一个像素的下一个通道,根本到不了下一个像素
  • 第2维对应通道:从像素的一个通道移动到下一个通道,只需要跨过1个uint8_t的长度,步长应为sizeof(uint8_t),你写的步长是3,直接跨到了3个元素之后

另外你直接把栈上局部变量cvimg3的data指针交给numpy数组包装,函数返回时cvimg3会被自动析构,对应内存被释放,Python端访问时会触发野指针,轻则读到随机脏数据,重则直接程序崩溃。

修复方法

简单稳妥写法(推荐)

直接让pybind11为返回的numpy数组分配独立内存,再把OpenCV Mat的数据拷贝过去,不用手动管理生命周期,不会踩内存坑:

py::array_t<uint8_t> flipcvMat(py::array_t<uint8_t>& img)
{
    auto rows = img.shape(0);
    auto cols = img.shape(1);
    auto channels = img.shape(2);
    auto type = CV_MAKETYPE(CV_8U, channels); // 不要硬编码CV_8UC3,适配任意通道数输入

    cv::Mat cvimg2(rows, cols, type, (unsigned char*)img.data());
    cv::Mat cvimg3(rows, cols, type);
    cv::flip(cvimg2, cvimg3, 0);

    // 构造对应维度的numpy数组,自动计算正确步长、分配内存
    py::array_t<uint8_t> output({rows, cols, channels});
    auto buf_info = output.request();
    // 把处理完的Mat数据拷贝到numpy数组内存中
    memcpy(buf_info.ptr, cvimg3.data, rows * cols * channels * sizeof(uint8_t));
    
    return output;
}

零拷贝写法(性能优先)

如果图像分辨率大、不想做内存拷贝,可以通过pybind11的capsule机制绑定内存生命周期,保证OpenCV Mat的内存直到numpy数组被垃圾回收时才释放:

py::array_t<uint8_t> flipcvMat(py::array_t<uint8_t>& img)
{
    auto rows = img.shape(0);
    auto cols = img.shape(1);
    auto channels = img.shape(2);
    auto type = CV_MAKETYPE(CV_8U, channels);

    cv::Mat cvimg2(rows, cols, type, (unsigned char*)img.data());
    // 把处理后的Mat放到堆上,交给capsule管理释放
    cv::Mat* cvimg3 = new cv::Mat(rows, cols, type);
    cv::flip(cvimg2, *cvimg3, 0);

    // 计算正确步长
    size_t ch_step = sizeof(uint8_t);
    size_t col_step = channels * ch_step;
    size_t row_step = cols * col_step;

    // 定义内存释放逻辑:numpy数组回收时自动delete堆上的Mat
    py::capsule mat_holder(cvimg3, [](void* ptr) {
        delete static_cast<cv::Mat*>(ptr);
    });

    // 构造零拷贝的numpy数组
    py::array_t<uint8_t> output(
        py::buffer_info(
            cvimg3->data,
            sizeof(uint8_t),
            py::format_descriptor<uint8_t>::format(),
            3,
            {rows, cols, channels},
            {row_step, col_step, ch_step}
        ),
        mat_holder
    );

    return output;
}

内容的提问来源于stack exchange,提问作者Ivan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 16:01:07