You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

QtConcurrent异步队列操作导致QImage内存未及时释放问题

Qt6.5.3多线程图像处理应用内存占用异常问题

我基于Qt6.5.3开发了一款图像处理应用,包含负责图像采集的生产者(模拟相机)和执行检测的消费者。由于检测速度较慢,采用多线程加速流程,核心代码如下:

#include <QCoreApplication>
#include <QDebug>
#include <QImage>
#include <QThread>
#include <QTimer>
#include <QtConcurrent>

class Producer : public QObject {
  Q_OBJECT

public:
  Producer(QObject *parent = nullptr) : QObject(parent) {}

public slots:
  void produce() {
    constexpr auto count = 1000;
    for (int i = 0; i < count; ++i) {
      QImage img(2448, 2048, QImage::Format_Grayscale8);
      img.fill(0);
      emit imageReady(img);
    }
  }

signals:
  void imageReady(QImage image);
};

class Consumer : public QObject {
  Q_OBJECT

public:
  Consumer(QObject *parent = nullptr) : QObject(parent) {}
  int consumedCount() const { return count_; }

public slots:
  void onImageReady(QImage image) {
    QFuture<void> future = QtConcurrent::run([=] {
      QImage copy = image.copy(); // Make a deep copy first
      QThread::msleep(200);       // Mock detection on the copy
      qDebug() << ++count_;
    });
  }

private:
  std::atomic_int count_ = 0;
};

int main(int argc, char *argv[]) {
  QCoreApplication a(argc, argv);
  Producer producer;
  Consumer consumer;
  QObject::connect(&producer, &Producer::imageReady, &consumer,
                   &Consumer::onImageReady);
  QTimer::singleShot(0, &producer, &Producer::produce);
  return a.exec();
}

#include "main.moc"

现象描述

消费者处理速度远慢于生产者,运行时图像排队等待检测占用约5GB内存属于预期,但所有图像检测完成后,进程内存占用仍维持在2GB左右,未正常释放。

排查过程

  1. 最初怀疑内存泄漏,但Valgrind Memcheck检测排除了该可能:
==123984== Memcheck, a memory error detector
==123984== Copyright (C) 2002-2017, and GNU GPL'd, by Julian Seward et al.
==123984== Using Valgrind-3.18.1 and LibVEX; rerun with -h for copyright info
==123984== Command: ./MultithreadImage
==123984== Parent PID: 99641
==123984== 
==123984== 
==123984== Process terminating with default action of signal 2 (SIGINT)
==123984==    at 0x5CA9BCF: poll (poll.c:29)
==123984==    by 0x60091F5: ??? (in /usr/lib/x86_64-linux-gnu/libglib-2.0.so.0.7200.4)
==123984==    by 0x5FB13E2: g_main_context_iteration (in /usr/lib/x86_64-linux-gnu/libglib-2.0.so.0.7200.4)
==123984==    by 0x569B809: QEventDispatcherGlib::processEvents(QFlags<QEventLoop::ProcessEventsFlag>) (qeventdispatcher_glib.cpp:393)
==123984==    by 0x53FCF6A: QEventLoop::exec(QFlags<QEventLoop::ProcessEventsFlag>) (qeventloop.cpp:182)
==123984==    by 0x53F97CD: QCoreApplication::exec() (qcoreapplication.cpp:1439)
==123984==    by 0x10B94F: main (main.cpp:55)
==123984== 
==123984== HEAP SUMMARY:
==123984==     in use at exit: 187,502 bytes in 332 blocks
==123984==   total heap usage: 16,974 allocs, 16,642 frees, 10,028,770,623 bytes allocated
==123984== 
==123984== LEAK SUMMARY:
==123984==    definitely lost: 0 bytes in 0 blocks
==123984==    indirectly lost: 0 bytes in 0 blocks
==123984==      possibly lost: 1,648 bytes in 7 blocks
==123984==    still reachable: 185,854 bytes in 325 blocks
==123984==                       of which reachable via heuristic:
==123984==                         newarray           : 328 bytes in 3 blocks
==123984==         suppressed: 0 bytes in 0 blocks
==123984== Rerun with --leak-check=full to see details of leaked memory
==123984== 
==123984== For lists of detected and suppressed errors, rerun with: -s
==123984== ERROR SUMMARY: 0 errors from 0 contexts (suppressed: 0 from 0)
  1. 调试发现QImage内部的引用计数直到应用即将退出时才归0,说明内存未泄漏,只是被应用内部持有。
  2. 注释掉QThread::msleep(200);让消费者与生产者近乎同步时,内存占用恢复正常,推测问题出在图像排队阶段。

问题原因

  1. 事件队列堆积导致QImage长期持有:生产者在主线程的produce函数中一次性发送1000个imageReady信号,此时主线程事件循环被produce函数阻塞,无法立即处理槽函数,所有信号参数(QImage,隐式共享的浅拷贝)会被堆积在事件队列中,直到produce执行完毕才开始处理。这些队列中的QImage会持续占用内存,直到对应的槽函数被调用。
  2. Qt内存分配器缓存机制:当所有QImage的引用计数归0后,Qt的内存分配器不会立即将释放的内存归还给操作系统,而是将其缓存起来,用于后续内存分配,避免频繁系统调用开销。这会导致进程的常驻内存(RSS)数值不会立即下降,看起来像是内存未释放。
  3. 异步任务持有QImage引用:槽函数中通过QtConcurrent::run启动的异步任务捕获了QImage的浅拷贝,该引用会持续到任务执行完毕,进一步延长了内存持有时间。

解决方法

  1. 控制生产者速度,避免事件队列堆积:通过信号量或自定义队列实现生产者-消费者的流量控制,比如当消费者处理的任务数低于阈值时,生产者才继续生成图像,防止事件队列堆积过多QImage。示例代码(简化版):
// 在Producer类中添加信号量成员
QSemaphore semaphore(10); // 限制队列最多10个待处理图像

void produce() {
    constexpr auto count = 1000;
    for (int i = 0; i < count; ++i) {
        semaphore.acquire(); // 等待信号量
        QImage img(2448, 2048, QImage::Format_Grayscale8);
        img.fill(0);
        emit imageReady(img);
    }
}

// 在Consumer的异步任务完成后释放信号量
void onImageReady(QImage image) {
    QImage copy = image.copy();
    QtConcurrent::run([this, copy = std::move(copy)]() mutable {
        QThread::msleep(200);
        qDebug() << ++count_;
        semaphore.release(); // 释放信号量,允许生产者继续生成
    });
}
  1. 优化QImage传递,减少引用持有时间:在槽函数中立即对QImage进行深拷贝,并通过移动语义将深拷贝的对象传递给异步任务,让原始的浅拷贝QImage可以尽早销毁:
void onImageReady(QImage image) {
    QImage copy = image.copy(); // 立即深拷贝
    QtConcurrent::run([this, copy = std::move(copy)]() mutable {
        QThread::msleep(200);
        qDebug() << ++count_;
    });
}
  1. 手动触发内存回收(可选):在所有任务完成后,调用系统或Qt的内存回收接口,将缓存的内存归还给操作系统。比如Linux下可以调用:
#include <malloc.h>
// 所有任务完成后执行
malloc_trim(0);

Windows下可以调用:

#include <windows.h>
// 所有任务完成后执行
HeapCompact(GetProcessHeap(), 0);

内容的提问来源于stack exchange,提问作者tanjor

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 01:37:31