Mesh优化咨询:网格生成类顶点数量过高导致帧率骤降,已用OptimizeInplace仅提升10帧求进阶方案
It’s frustrating when frame rates drop sharply with higher vertex counts, especially after already using OptimizeInplace. Let’s go through several targeted optimizations that can give you much bigger gains than the 10 FPS you’re seeing now, with specific changes to your code where possible.
1. Move Your Mesh to Video Memory (Critical!)
Right now, you’re loading your mesh into D3DXMESH_SYSTEMMEM, which stores data in CPU system memory. Every frame, the GPU has to pull this data across the bus, which is a massive bottleneck for high-vertex meshes.
Switch to either D3DXMESH_MANAGED (lets Direct3D handle moving data between system and video memory as needed) or D3DXMESH_WRITEONLY (for static meshes you won’t modify after loading—this is the fastest option since it lives exclusively in video memory). Adding D3DXMESH_32BIT also avoids index overflow issues for meshes with more than 65535 vertices.
Code Change:
Modify your D3DXLoadMeshFromX call:
D3DXLoadMeshFromX(fileName.c_str(), D3DXMESH_WRITEONLY | D3DXMESH_32BIT, Render->g_pDevice, NULL, &pD3DXMtrlBuffer, NULL, &m_dwNumMaterials, &m_pMesh);
2. Enhance Mesh Optimization Flags
Your current OptimizeInplace flags are a good start, but adding D3DXMESHOPT_STRIPREORDER will reorder triangles into strips—this reduces the number of vertices the GPU needs to process (triangle strips reuse the last two vertices for each new triangle instead of sending all three).
To specify a larger vertex cache size (most modern GPUs support 32 or 64, vs. the default 16), use D3DXOptimizeMesh instead of OptimizeInplace:
Code Change:
// Replace your existing OptimizeInplace call with this: LPD3DXMESH pOptMesh = NULL; D3DXOptimizeMesh(m_pMesh, &adjacencyBuffer[0], D3DXMESHOPT_ATTRSORT | D3DXMESHOPT_COMPACT | D3DXMESHOPT_VERTEXCACHE | D3DXMESHOPT_STRIPREORDER, NULL, 32, // Adjust based on your GPU's cache size &pOptMesh); // Swap to the optimized mesh m_pMesh->Release(); m_pMesh = pOptMesh;
3. Trim Your Vertex Format
If your mesh includes unused vertex data (e.g., extra UV channels, tangents/binormals you don’t need for rendering), this wastes memory bandwidth and GPU processing time. Simplify your vertex format to only include what’s necessary.
For example, if you don’t use normal mapping, convert your mesh to remove tangents/binormals:
LPD3DXMESH pSimplifiedMesh = NULL; D3DXConvertMeshFVF(m_pMesh, D3DFVF_XYZ | D3DFVF_NORMAL | D3DFVF_TEX1, // Keep only position, normal, single UV set Render->g_pDevice, &pSimplifiedMesh); m_pMesh->Release(); m_pMesh = pSimplifiedMesh;
4. Compress Textures to Reduce Bandwidth
Your current D3DXCreateTextureFromFileA loads textures in their uncompressed original format, which is memory-heavy. Compressed formats like DXT1 (opaque) or DXT5 (transparent) cut memory usage by 75% and drastically reduce GPU bandwidth overhead.
Code Change:
if (d3dxMaterials[i].pTextureFilename != NULL) { D3DXCreateTextureFromFileExA( Render->g_pDevice, d3dxMaterials[i].pTextureFilename, D3DX_DEFAULT_NONPOW2, // Allow non-power-of-two if your GPU supports it D3DX_DEFAULT_NONPOW2, D3DX_DEFAULT, 0, D3DFMT_DXT5, // Use D3DFMT_DXT1 for opaque textures D3DPOOL_MANAGED, D3DX_FILTER_TRIANGLE | D3DX_FILTER_MIRROR, D3DX_FILTER_TRIANGLE | D3DX_FILTER_MIRROR, 0, NULL, NULL, &m_pMeshTextures[i] ); }
5. Batch Draw Calls to Minimize State Changes
Your render function calls DrawSubset for each material, which involves repeated expensive state changes (setting materials/textures). If multiple subsets share the same material/texture, batch these draws together:
Implementation Idea:
First, build a map of materials to subset indices in your initialize method:
// Add this after loading materials/textures std::unordered_map<LPDIRECT3DTEXTURE9, std::vector<DWORD>> materialSubsets; for (DWORD i = 0; i < m_dwNumMaterials; i++) { materialSubsets[m_pMeshTextures[i]].push_back(i); }
Then update your render method to batch draws:
void Sprite::render() { Render->g_pDevice->SetRenderState(D3DRS_FILLMODE, D3DFILL_SOLID); for (auto& materialGroup : materialSubsets) { // Set state once per material Render->g_pDevice->SetMaterial(&m_pMeshMaterials[materialGroup.second[0]]); Render->g_pDevice->SetTexture(0, materialGroup.first); // Draw all subsets for this material for (DWORD subset : materialGroup.second) { m_pMesh->DrawSubset(subset); } } }
6. Use Hardware Instancing (If Rendering Multiple Instances)
If you’re drawing multiple copies of the same Sprite mesh, hardware instancing lets you render all instances with a single draw call. This eliminates the overhead of repeated DrawSubset calls and is a massive performance win for multiple instances. You’ll need to add instance data (like world matrices) to your mesh and use DrawIndexedPrimitiveUP with instancing support.
Combining these optimizations—especially moving the mesh to video memory and adding strip reordering—should deliver significant frame rate improvements even with high vertex counts.
内容的提问来源于stack exchange,提问作者Meggii

