You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C++中高效向PyList追加图像像素值的优化方案

优化C++向PyList追加图像像素的效率

我是C++新手,当前处理单帧854×480分辨率图像时,将像素值逐个追加到PyList的操作耗时约0.1秒,希望找到更高效的实现方法,尽量避免使用第三方库。以下是我当前的实现代码:

PyObject* byte_list = PyList_New(static_cast<Py_ssize_t>(0));

AVFrame *pFrameRGB = av_frame_alloc();
av_frame_copy_props(pFrameRGB, this->pFrame);
pFrameRGB->width = this->pFrame->width;
pFrameRGB->height = this->pFrame->height;
pFrameRGB->format = AV_PIX_FMT_RGB24;
av_frame_get_buffer(pFrameRGB, 0);

sws_scale(this->swsCtx, this->pFrame->data, this->pFrame->linesize, 0, 
        this->pCodecContext->height, pFrameRGB->data, pFrameRGB->linesize);

if (this->_debug) {
    std::cout << "Frame linesize " << pFrameRGB->linesize[0] << "\n";
    std::cout << "Frame width " << pFrameRGB->width << "\n";
    std::cout << "Frame height " << pFrameRGB->height << "\n";
}

// 这段循环速度很慢
for(int y = 0; y < pFrameRGB->height; ++y) {
    for(int x = 0; x < pFrameRGB->width; ++x) {
        int p = x * 3 + y * pFrameRGB->linesize[0];
        int r = pFrameRGB->data[0][p];
        int g = pFrameRGB->data[0][p+1];
        int b = pFrameRGB->data[0][p+2];
        PyList_Append(byte_list, PyLong_FromLong(r));
        PyList_Append(byte_list, PyLong_FromLong(g));
        PyList_Append(byte_list, PyLong_FromLong(b));
    }
}

av_frame_free(&pFrameRGB);

优化方案及代码实现

核心性能瓶颈分析

原代码的主要耗时点在于:

  • PyList_Append会频繁触发列表扩容,每次扩容都要重新分配内存并拷贝数据。
  • 每次调用PyLong_FromLong都会创建新的Python整数对象,带来大量的对象构造和销毁开销。
  • 逐像素循环的内存访问模式不够高效。

优化方案一:预分配列表+复用Python整数对象

// 提前预创建0-255的PyLong对象缓存,避免重复创建
static PyObject* byte_cache[256] = {nullptr};
if (!byte_cache[0]) {
    for (int i = 0; i < 256; ++i) {
        byte_cache[i] = PyLong_FromLong(i);
        Py_INCREF(byte_cache[i]); // 增加引用计数,防止被Python垃圾回收
    }
}

// 预分配足够容量的PyList:总元素数 = 宽度 × 高度 × 3(RGB三通道)
Py_ssize_t total_items = static_cast<Py_ssize_t>(pFrameRGB->width * pFrameRGB->height * 3);
PyObject* byte_list = PyList_New(total_items);

Py_ssize_t list_idx = 0;
for(int y = 0; y < pFrameRGB->height; ++y) {
    // 直接获取当前行的起始指针,减少计算量
    uint8_t* row_start = pFrameRGB->data[0] + y * pFrameRGB->linesize[0];
    for(int x = 0; x < pFrameRGB->width; ++x) {
        uint8_t r = row_start[x*3];
        uint8_t g = row_start[x*3 + 1];
        uint8_t b = row_start[x*3 + 2];
        // 直接设置列表元素,替代低效的PyList_Append
        PyList_SetItem(byte_list, list_idx++, byte_cache[r]);
        PyList_SetItem(byte_list, list_idx++, byte_cache[g]);
        PyList_SetItem(byte_list, list_idx++, byte_cache[b]);
    }
}

优化方案二:利用Python bytes对象批量转换(效率更高)

如果业务允许,先将RGB数据打包成Python的bytes对象,再转成列表,Python内部会做高效的批量处理:

// 计算有效RGB数据总字节数
size_t total_rgb_bytes = pFrameRGB->width * pFrameRGB->height * 3;
// 开辟临时缓冲区,拷贝每行的有效像素(跳过linesize的padding)
std::vector<uint8_t> rgb_buffer(total_rgb_bytes);
uint8_t* dst_ptr = rgb_buffer.data();

for(int y = 0; y < pFrameRGB->height; ++y) {
    uint8_t* row_start = pFrameRGB->data[0] + y * pFrameRGB->linesize[0];
    memcpy(dst_ptr, row_start, pFrameRGB->width * 3);
    dst_ptr += pFrameRGB->width * 3;
}

// 创建Python bytes对象
PyObject* rgb_bytes = PyBytes_FromStringAndSize(reinterpret_cast<const char*>(rgb_buffer.data()), total_rgb_bytes);
// 将bytes转换为列表,内部批量处理远快于逐元素添加
PyObject* byte_list = PySequence_List(rgb_bytes);

// 释放临时对象
Py_DECREF(rgb_bytes);

内容的提问来源于stack exchange,提问作者Timothy Halim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 12:45:39