C++中高效向PyList追加图像像素值的优化方案
优化C++向PyList追加图像像素的效率
我是C++新手,当前处理单帧854×480分辨率图像时,将像素值逐个追加到PyList的操作耗时约0.1秒,希望找到更高效的实现方法,尽量避免使用第三方库。以下是我当前的实现代码:
PyObject* byte_list = PyList_New(static_cast<Py_ssize_t>(0)); AVFrame *pFrameRGB = av_frame_alloc(); av_frame_copy_props(pFrameRGB, this->pFrame); pFrameRGB->width = this->pFrame->width; pFrameRGB->height = this->pFrame->height; pFrameRGB->format = AV_PIX_FMT_RGB24; av_frame_get_buffer(pFrameRGB, 0); sws_scale(this->swsCtx, this->pFrame->data, this->pFrame->linesize, 0, this->pCodecContext->height, pFrameRGB->data, pFrameRGB->linesize); if (this->_debug) { std::cout << "Frame linesize " << pFrameRGB->linesize[0] << "\n"; std::cout << "Frame width " << pFrameRGB->width << "\n"; std::cout << "Frame height " << pFrameRGB->height << "\n"; } // 这段循环速度很慢 for(int y = 0; y < pFrameRGB->height; ++y) { for(int x = 0; x < pFrameRGB->width; ++x) { int p = x * 3 + y * pFrameRGB->linesize[0]; int r = pFrameRGB->data[0][p]; int g = pFrameRGB->data[0][p+1]; int b = pFrameRGB->data[0][p+2]; PyList_Append(byte_list, PyLong_FromLong(r)); PyList_Append(byte_list, PyLong_FromLong(g)); PyList_Append(byte_list, PyLong_FromLong(b)); } } av_frame_free(&pFrameRGB);
优化方案及代码实现
核心性能瓶颈分析
原代码的主要耗时点在于:
PyList_Append会频繁触发列表扩容,每次扩容都要重新分配内存并拷贝数据。- 每次调用
PyLong_FromLong都会创建新的Python整数对象,带来大量的对象构造和销毁开销。 - 逐像素循环的内存访问模式不够高效。
优化方案一:预分配列表+复用Python整数对象
// 提前预创建0-255的PyLong对象缓存,避免重复创建 static PyObject* byte_cache[256] = {nullptr}; if (!byte_cache[0]) { for (int i = 0; i < 256; ++i) { byte_cache[i] = PyLong_FromLong(i); Py_INCREF(byte_cache[i]); // 增加引用计数,防止被Python垃圾回收 } } // 预分配足够容量的PyList:总元素数 = 宽度 × 高度 × 3(RGB三通道) Py_ssize_t total_items = static_cast<Py_ssize_t>(pFrameRGB->width * pFrameRGB->height * 3); PyObject* byte_list = PyList_New(total_items); Py_ssize_t list_idx = 0; for(int y = 0; y < pFrameRGB->height; ++y) { // 直接获取当前行的起始指针,减少计算量 uint8_t* row_start = pFrameRGB->data[0] + y * pFrameRGB->linesize[0]; for(int x = 0; x < pFrameRGB->width; ++x) { uint8_t r = row_start[x*3]; uint8_t g = row_start[x*3 + 1]; uint8_t b = row_start[x*3 + 2]; // 直接设置列表元素,替代低效的PyList_Append PyList_SetItem(byte_list, list_idx++, byte_cache[r]); PyList_SetItem(byte_list, list_idx++, byte_cache[g]); PyList_SetItem(byte_list, list_idx++, byte_cache[b]); } }
优化方案二:利用Python bytes对象批量转换(效率更高)
如果业务允许,先将RGB数据打包成Python的bytes对象,再转成列表,Python内部会做高效的批量处理:
// 计算有效RGB数据总字节数 size_t total_rgb_bytes = pFrameRGB->width * pFrameRGB->height * 3; // 开辟临时缓冲区,拷贝每行的有效像素(跳过linesize的padding) std::vector<uint8_t> rgb_buffer(total_rgb_bytes); uint8_t* dst_ptr = rgb_buffer.data(); for(int y = 0; y < pFrameRGB->height; ++y) { uint8_t* row_start = pFrameRGB->data[0] + y * pFrameRGB->linesize[0]; memcpy(dst_ptr, row_start, pFrameRGB->width * 3); dst_ptr += pFrameRGB->width * 3; } // 创建Python bytes对象 PyObject* rgb_bytes = PyBytes_FromStringAndSize(reinterpret_cast<const char*>(rgb_buffer.data()), total_rgb_bytes); // 将bytes转换为列表,内部批量处理远快于逐元素添加 PyObject* byte_list = PySequence_List(rgb_bytes); // 释放临时对象 Py_DECREF(rgb_bytes);
内容的提问来源于stack exchange,提问作者Timothy Halim
相关产品推荐
相关产品推荐

