Pybind11绑定C++ cv::Mat返回Python端图像像素错位问题
问题描述
我需要编写一个接收图像作为入参、返回处理后图像的函数,使用pybind11将该函数绑定到Python环境调用。
目前Python端传图像到C的逻辑已经跑通,但C处理完的图像返回Python环节出现异常。
以下是实现图像翻转功能的示例代码:
py::array_t<uint8_t> flipcvMat(py::array_t<uint8_t>& img) { auto rows = img.shape(0); auto cols = img.shape(1); auto channels = img.shape(2); std::cout << "rows: " << rows << " cols: " << cols << " channels: " << channels << std::endl; auto type = CV_8UC3; cv::Mat cvimg2(rows, cols, type, (unsigned char*)img.data()); cv::imwrite("/source/test.png", cvimg2); // 本地写入正常 cv::Mat cvimg3(rows, cols, type); cv::flip(cvimg2, cvimg3, 0); cv::imwrite("/source/testout.png", cvimg3); // 本地写入正常 py::array_t<uint8_t> output( py::buffer_info( cvimg3.data, sizeof(uint8_t), //itemsize py::format_descriptor<uint8_t>::format(), 3, // ndim std::vector<size_t> {rows, cols , 3}, // shape std::vector<size_t> {cols * sizeof(uint8_t), sizeof(uint8_t), 3} // strides ) ); return output; }
Python端调用代码:
import cv2 img = cv2.imread('/source/whatever/ubuntu-1.png') img3= opencvtest.flipcvMat(img)
已知验证结果:
- 输入为3通道图像
- C++侧通过
cv::imwrite保存的输入副本、翻转后结果图像均正常,翻转逻辑本身没有问题
故障现象:Python端接收到的返回图像存在像素错位,只能模糊识别轮廓,像素排列完全混乱。初步判断是构造py::buffer_info时参数配置错误,但未定位到具体问题。
问题根因
核心错误是py::buffer_info中strides步长参数配置错误,另外还隐藏了内存生命周期的野指针风险。
步长的定义是:从当前维度的一个元素移动到同维度下一个相邻元素,需要跨过的字节总数。你当前写的步长{cols * sizeof(uint8_t), sizeof(uint8_t), 3}完全不符合HWC格式三通道图像的内存排布规则:
- 第0维对应图像行(高):从一行首像素移动到下一行首像素,需要跨过整行所有像素的所有通道,步长应为
cols * channels * sizeof(uint8_t),你写的步长没有乘通道数,相当于每读一行只跳了1/3的行长度,必然错位 - 第1维对应图像列(宽):从某列像素移动到下一列同位置像素,需要跨过1个像素的全部通道,步长应为
channels * sizeof(uint8_t),你写的步长只跳1字节,每次只挪到同一个像素的下一个通道,根本到不了下一个像素 - 第2维对应通道:从像素的一个通道移动到下一个通道,只需要跨过1个uint8_t的长度,步长应为
sizeof(uint8_t),你写的步长是3,直接跨到了3个元素之后
另外你直接把栈上局部变量cvimg3的data指针交给numpy数组包装,函数返回时cvimg3会被自动析构,对应内存被释放,Python端访问时会触发野指针,轻则读到随机脏数据,重则直接程序崩溃。
修复方法
简单稳妥写法(推荐)
直接让pybind11为返回的numpy数组分配独立内存,再把OpenCV Mat的数据拷贝过去,不用手动管理生命周期,不会踩内存坑:
py::array_t<uint8_t> flipcvMat(py::array_t<uint8_t>& img) { auto rows = img.shape(0); auto cols = img.shape(1); auto channels = img.shape(2); auto type = CV_MAKETYPE(CV_8U, channels); // 不要硬编码CV_8UC3,适配任意通道数输入 cv::Mat cvimg2(rows, cols, type, (unsigned char*)img.data()); cv::Mat cvimg3(rows, cols, type); cv::flip(cvimg2, cvimg3, 0); // 构造对应维度的numpy数组,自动计算正确步长、分配内存 py::array_t<uint8_t> output({rows, cols, channels}); auto buf_info = output.request(); // 把处理完的Mat数据拷贝到numpy数组内存中 memcpy(buf_info.ptr, cvimg3.data, rows * cols * channels * sizeof(uint8_t)); return output; }
零拷贝写法(性能优先)
如果图像分辨率大、不想做内存拷贝,可以通过pybind11的capsule机制绑定内存生命周期,保证OpenCV Mat的内存直到numpy数组被垃圾回收时才释放:
py::array_t<uint8_t> flipcvMat(py::array_t<uint8_t>& img) { auto rows = img.shape(0); auto cols = img.shape(1); auto channels = img.shape(2); auto type = CV_MAKETYPE(CV_8U, channels); cv::Mat cvimg2(rows, cols, type, (unsigned char*)img.data()); // 把处理后的Mat放到堆上,交给capsule管理释放 cv::Mat* cvimg3 = new cv::Mat(rows, cols, type); cv::flip(cvimg2, *cvimg3, 0); // 计算正确步长 size_t ch_step = sizeof(uint8_t); size_t col_step = channels * ch_step; size_t row_step = cols * col_step; // 定义内存释放逻辑:numpy数组回收时自动delete堆上的Mat py::capsule mat_holder(cvimg3, [](void* ptr) { delete static_cast<cv::Mat*>(ptr); }); // 构造零拷贝的numpy数组 py::array_t<uint8_t> output( py::buffer_info( cvimg3->data, sizeof(uint8_t), py::format_descriptor<uint8_t>::format(), 3, {rows, cols, channels}, {row_step, col_step, ch_step} ), mat_holder ); return output; }
内容的提问来源于stack exchange,提问作者Ivan
相关产品推荐
相关产品推荐

