如何提升QQuickItem渲染性能?20万三角面帧率过低优化求助
问题
我正在开发一款渲染大量三角面的QML应用,目前尝试渲染200000个三角面。我的显卡是RTX 2060,理论上可在60fps下渲染数百万三角面,但应用仅能达到约10fps。经排查确定为GPU瓶颈:CPU使用率极低,而GPU始终处于100%占用状态。
我曾尝试通过QCoreApplication::setAttribute(Qt::AA_UseOpenGLES);切换OpenGL后端,但无效果。请问我该如何提升性能?
代码实现
main.cpp
#include <QGuiApplication> #include <QQmlApplicationEngine> #include <QQuickWindow> #include "DenemeClass.h" int main(int argc, char *argv[]) { #if QT_VERSION < QT_VERSION_CHECK(6, 0, 0) QCoreApplication::setAttribute(Qt::AA_EnableHighDpiScaling); #endif QGuiApplication app(argc, argv); qmlRegisterType<DenemeClass>("com.deneme", 1, 0, "DenemeClass"); QQmlApplicationEngine engine; const QUrl url(QStringLiteral("qrc:/main.qml")); QObject::connect(&engine, &QQmlApplicationEngine::objectCreated, &app, [url](QObject *obj, const QUrl &objUrl) { if (!obj && url == objUrl) QCoreApplication::exit(-1); }, Qt::QueuedConnection); engine.load(url); return app.exec(); }
DenemeClass.h
#ifndef DENEMECLASS_H #define DENEMECLASS_H #include <QQuickItem> class DenemeClass : public QQuickItem { Q_OBJECT public: DenemeClass(QQuickItem* parent = nullptr); QSGNode* updatePaintNode(QSGNode* oldNode, UpdatePaintNodeData* data) override; int64_t lastDraw; }; #endif // DENEMECLASS_H
DenemeClass.cpp
#include "DenemeClass.h" #include <QSGGeometryNode> #include <QSGFlatColorMaterial> #include <QSGGeometry> #include <iostream> DenemeClass::DenemeClass(QQuickItem* parent) : QQuickItem(parent) { setFlag(ItemHasContents, true); } QSGNode* DenemeClass::updatePaintNode(QSGNode* oldNode, UpdatePaintNodeData* data){ int64_t draw_t = std::chrono::duration_cast<std::chrono::milliseconds>(std::chrono::steady_clock::now().time_since_epoch()).count(); std::cout << (1000.0 / double(draw_t - lastDraw)) << "FPS" << std::endl; lastDraw = draw_t; const int rectCount = 100000; QSGGeometryNode* node = reinterpret_cast<QSGGeometryNode*>(oldNode); if(!node){ node = new QSGGeometryNode(); node->setFlag(QSGNode::OwnsMaterial, true); node->setFlag(QSGNode::OwnsGeometry, true); QSGFlatColorMaterial* material = new QSGFlatColorMaterial; material->setColor("#ff0000"); node->setMaterial(material); QSGGeometry* geometry = new QSGGeometry(QSGGeometry::defaultAttributes_Point2D(), rectCount * 6, 0, QSGGeometry::UnsignedIntType); geometry->setDrawingMode(QSGGeometry::DrawTriangles); node->setGeometry(geometry); QSGGeometry::Point2D* pts = geometry->vertexDataAsPoint2D(); for(int i = 0; i < rectCount; ++i){ pts[i * 6 + 0].x = 0; pts[i * 6 + 0].y = 0; pts[i * 6 + 1].x = 0; pts[i * 6 + 1].y = height(); pts[i * 6 + 2].x = width(); pts[i * 6 + 2].y = height(); pts[i * 6 + 3].x = width(); pts[i * 6 + 3].y = height(); pts[i * 6 + 4].x = width(); pts[i * 6 + 4].y = 0; pts[i * 6 + 5].x = 0; pts[i * 6 + 5].y = 0; } } return node; }
main.qml
import QtQuick 2.15 import QtQuick.Window 2.15 import com.deneme 1.0 Window { width: 640 height: 480 visible: true title: qsTr("Hello World") property double xx: 0 NumberAnimation on xx{ running: true loops: Animation.Infinite from: 0 to: 100 duration: 2000 } Rectangle{ anchors.fill: parent color: "yellow" anchors.margins: xx } DenemeClass{ anchors.fill: parent anchors.margins: xx } }
优化方案
1. 使用索引缓冲减少顶点数据量
当前每个矩形用6个重复顶点绘制两个三角面,通过索引缓冲可复用顶点,将顶点总数从rectCount*6降至rectCount*4,仅需增加rectCount*6个索引。这能大幅减少GPU需要处理的顶点数据量,降低内存带宽占用。
修改代码示例:
// 创建几何时指定顶点数和索引数 QSGGeometry* geometry = new QSGGeometry(QSGGeometry::defaultAttributes_Point2D(), rectCount * 4, rectCount * 6); geometry->setDrawingMode(QSGGeometry::DrawTriangles); QSGGeometry::Point2D* pts = geometry->vertexDataAsPoint2D(); quint32* indices = geometry->indexDataAsUInt32(); for(int i = 0; i < rectCount; ++i){ int baseVert = i * 4; // 单个矩形的4个顶点 pts[baseVert + 0].x = 0; pts[baseVert + 0].y = 0; pts[baseVert + 1].x = 0; pts[baseVert + 1].y = height(); pts[baseVert + 2].x = width(); pts[baseVert + 2].y = height(); pts[baseVert + 3].x = width(); pts[baseVert + 3].y = 0; // 索引复用顶点绘制两个三角面 int baseIdx = i * 6; indices[baseIdx + 0] = baseVert + 0; indices[baseIdx + 1] = baseVert + 1; indices[baseIdx + 2] = baseVert + 2; indices[baseIdx + 3] = baseVert + 2; indices[baseIdx + 4] = baseVert + 3; indices[baseIdx + 5] = baseVert + 0; }
2. 启用实例化渲染
如果所有矩形的材质、渲染状态一致,使用实例化渲染可让GPU一次完成所有矩形的绘制,减少CPU到GPU的命令开销,提升GPU渲染效率。Qt QSG支持通过QSGGeometry::setInstanced(true)结合自定义顶点属性传递每个实例的位置、尺寸数据,实现单绘制调用渲染所有实例。
3. 用着色器处理变换,避免频繁更新顶点数据
当前Item尺寸随动画变化时,直接修改顶点数据会导致GPU重复上传数据。可将顶点设置为单位矩形(0,0到1,1),在自定义着色器中通过uniform变量接收Item的宽度、高度和偏移,由GPU完成坐标变换,无需修改顶点数据,大幅降低数据传输开销。
4. 优化渲染状态与材质
- 若无需遮挡检测,通过
node->setFlag(QSGNode::NoDepthTest, true)关闭深度测试,减少GPU深度计算开销。 - 确保使用高效的图形后端:Qt 5.14+可设置
QQuickWindow::setSceneGraphBackend(QSGRendererInterface::OpenGLRhi),Qt 6默认使用RHI后端,避免老旧的OpenGL后端。
5. 减少QML层级与重叠绘制
当前QML中有两个全屏Item,会触发两次全屏绘制。若DenemeClass已覆盖底层Rectangle,可移除Rectangle,或直接在DenemeClass中设置黄色背景,减少一次绘制调用,降低GPU填充率开销。
内容的提问来源于stack exchange,提问作者Mertcan Özdemir

