You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提升QQuickItem渲染性能?20万三角面帧率过低优化求助

问题

我正在开发一款渲染大量三角面的QML应用,目前尝试渲染200000个三角面。我的显卡是RTX 2060,理论上可在60fps下渲染数百万三角面,但应用仅能达到约10fps。经排查确定为GPU瓶颈:CPU使用率极低,而GPU始终处于100%占用状态。

我曾尝试通过QCoreApplication::setAttribute(Qt::AA_UseOpenGLES);切换OpenGL后端,但无效果。请问我该如何提升性能?

代码实现

main.cpp

#include <QGuiApplication>
#include <QQmlApplicationEngine>
#include <QQuickWindow>
#include "DenemeClass.h"

int main(int argc, char *argv[])
{
#if QT_VERSION < QT_VERSION_CHECK(6, 0, 0)
    QCoreApplication::setAttribute(Qt::AA_EnableHighDpiScaling);
#endif
    QGuiApplication app(argc, argv);

    qmlRegisterType<DenemeClass>("com.deneme", 1, 0, "DenemeClass");

    QQmlApplicationEngine engine;
    const QUrl url(QStringLiteral("qrc:/main.qml"));
    QObject::connect(&engine, &QQmlApplicationEngine::objectCreated,
        &app, [url](QObject *obj, const QUrl &objUrl) {
            if (!obj && url == objUrl)
                QCoreApplication::exit(-1);
        }, Qt::QueuedConnection);
    engine.load(url);

    return app.exec();
}

DenemeClass.h

#ifndef DENEMECLASS_H
#define DENEMECLASS_H
#include <QQuickItem>

class DenemeClass : public QQuickItem
{
    Q_OBJECT
public:
    DenemeClass(QQuickItem* parent = nullptr);
    QSGNode* updatePaintNode(QSGNode* oldNode, UpdatePaintNodeData* data) override;
    int64_t lastDraw;
};

#endif // DENEMECLASS_H

DenemeClass.cpp

#include "DenemeClass.h"
#include <QSGGeometryNode>
#include <QSGFlatColorMaterial>
#include <QSGGeometry>
#include <iostream>
DenemeClass::DenemeClass(QQuickItem* parent)
    : QQuickItem(parent)
{
    setFlag(ItemHasContents, true);
}

QSGNode* DenemeClass::updatePaintNode(QSGNode* oldNode, UpdatePaintNodeData* data){

    int64_t draw_t = std::chrono::duration_cast<std::chrono::milliseconds>(std::chrono::steady_clock::now().time_since_epoch()).count();
    std::cout << (1000.0 / double(draw_t - lastDraw)) << "FPS" << std::endl;
    lastDraw = draw_t;

    const int rectCount = 100000;

    QSGGeometryNode* node = reinterpret_cast<QSGGeometryNode*>(oldNode);
    if(!node){
        node = new QSGGeometryNode();
        node->setFlag(QSGNode::OwnsMaterial, true);
        node->setFlag(QSGNode::OwnsGeometry, true);
        QSGFlatColorMaterial* material = new QSGFlatColorMaterial;
        material->setColor("#ff0000");
        node->setMaterial(material);

        QSGGeometry* geometry = new QSGGeometry(QSGGeometry::defaultAttributes_Point2D(), rectCount * 6, 0, QSGGeometry::UnsignedIntType);
        geometry->setDrawingMode(QSGGeometry::DrawTriangles);
        node->setGeometry(geometry);


        QSGGeometry::Point2D* pts = geometry->vertexDataAsPoint2D();

        for(int i = 0; i < rectCount; ++i){
            pts[i * 6 + 0].x = 0;
            pts[i * 6 + 0].y = 0;

            pts[i * 6 + 1].x = 0;
            pts[i * 6 + 1].y = height();

            pts[i * 6 + 2].x = width();
            pts[i * 6 + 2].y = height();

            pts[i * 6 + 3].x = width();
            pts[i * 6 + 3].y = height();

            pts[i * 6 + 4].x = width();
            pts[i * 6 + 4].y = 0;

            pts[i * 6 + 5].x = 0;
            pts[i * 6 + 5].y = 0;

        }
    }
    return node;
}

main.qml

import QtQuick 2.15
import QtQuick.Window 2.15
import com.deneme 1.0
Window {
    width: 640
    height: 480
    visible: true
    title: qsTr("Hello World")

    property double xx: 0
    NumberAnimation on xx{
        running: true
        loops: Animation.Infinite
        from: 0
        to: 100
        duration: 2000
    }

    Rectangle{
        anchors.fill: parent
        color: "yellow"
        anchors.margins: xx
    }

    DenemeClass{
        anchors.fill: parent
        anchors.margins: xx
    }
}

优化方案

1. 使用索引缓冲减少顶点数据量

当前每个矩形用6个重复顶点绘制两个三角面,通过索引缓冲可复用顶点,将顶点总数从rectCount*6降至rectCount*4,仅需增加rectCount*6个索引。这能大幅减少GPU需要处理的顶点数据量,降低内存带宽占用。

修改代码示例:

// 创建几何时指定顶点数和索引数
QSGGeometry* geometry = new QSGGeometry(QSGGeometry::defaultAttributes_Point2D(), rectCount * 4, rectCount * 6);
geometry->setDrawingMode(QSGGeometry::DrawTriangles);

QSGGeometry::Point2D* pts = geometry->vertexDataAsPoint2D();
quint32* indices = geometry->indexDataAsUInt32();

for(int i = 0; i < rectCount; ++i){
    int baseVert = i * 4;
    // 单个矩形的4个顶点
    pts[baseVert + 0].x = 0;
    pts[baseVert + 0].y = 0;
    pts[baseVert + 1].x = 0;
    pts[baseVert + 1].y = height();
    pts[baseVert + 2].x = width();
    pts[baseVert + 2].y = height();
    pts[baseVert + 3].x = width();
    pts[baseVert + 3].y = 0;

    // 索引复用顶点绘制两个三角面
    int baseIdx = i * 6;
    indices[baseIdx + 0] = baseVert + 0;
    indices[baseIdx + 1] = baseVert + 1;
    indices[baseIdx + 2] = baseVert + 2;
    indices[baseIdx + 3] = baseVert + 2;
    indices[baseIdx + 4] = baseVert + 3;
    indices[baseIdx + 5] = baseVert + 0;
}

2. 启用实例化渲染

如果所有矩形的材质、渲染状态一致,使用实例化渲染可让GPU一次完成所有矩形的绘制,减少CPU到GPU的命令开销,提升GPU渲染效率。Qt QSG支持通过QSGGeometry::setInstanced(true)结合自定义顶点属性传递每个实例的位置、尺寸数据,实现单绘制调用渲染所有实例。

3. 用着色器处理变换,避免频繁更新顶点数据

当前Item尺寸随动画变化时,直接修改顶点数据会导致GPU重复上传数据。可将顶点设置为单位矩形(0,0到1,1),在自定义着色器中通过uniform变量接收Item的宽度、高度和偏移,由GPU完成坐标变换,无需修改顶点数据,大幅降低数据传输开销。

4. 优化渲染状态与材质

  • 若无需遮挡检测,通过node->setFlag(QSGNode::NoDepthTest, true)关闭深度测试,减少GPU深度计算开销。
  • 确保使用高效的图形后端:Qt 5.14+可设置QQuickWindow::setSceneGraphBackend(QSGRendererInterface::OpenGLRhi),Qt 6默认使用RHI后端,避免老旧的OpenGL后端。

5. 减少QML层级与重叠绘制

当前QML中有两个全屏Item,会触发两次全屏绘制。若DenemeClass已覆盖底层Rectangle,可移除Rectangle,或直接在DenemeClass中设置黄色背景,减少一次绘制调用,降低GPU填充率开销。

内容的提问来源于stack exchange,提问作者Mertcan Özdemir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 19:28:11