C#中操作小型Bitmap像素的最优方法仍卡顿,如何提升性能?
Bitmap像素处理性能优化方案
以下是针对你的代码的具体优化措施,直接解决性能瓶颈:
1. 减少内存读写次数(核心优化)
当前代码每个Shader处理完就立即写回内存,造成大量重复的内存IO(这是性能杀手)。应该先一次性读出像素值,经过所有Shader处理后再写回内存:
{ BitmapData bitmapData = bmp.LockBits(new Rectangle(0, 0, bmp.Width, bmp.Height), ImageLockMode.ReadWrite, bmp.PixelFormat); int bytesPerPixel = Image.GetPixelFormatSize(bmp.PixelFormat) / 8; int heightInPixels = bitmapData.Height; int widthInBytes = bitmapData.Width * bytesPerPixel; byte* PtrFirstPixel = (byte*)bitmapData.Scan0; for(int y = 0; y < heightInPixels; y++) { byte* currentLine = PtrFirstPixel + (y * bitmapData.Stride); for (int x = 0; x < widthInBytes; x += bytesPerPixel) { // 一次性读出像素值 rgba color = new rgba(currentLine[x], currentLine[x + 1], currentLine[x + 2], -1); if (bytesPerPixel == 4) color.a = currentLine[x + 3]; // 依次执行所有Shader处理 for (int i = 0; i < Shaders.Count; i++) { // 修复原坐标计算错误:x/bytesPerPixel才是横坐标,纵坐标直接用y Shaders[i].pixel(new xy(x / bytesPerPixel, y), color); } // 最后一次性写回内存 currentLine[x] = (byte)color.r; currentLine[x + 1] = (byte)color.g; currentLine[x + 2] = (byte)color.b; if (bytesPerPixel == 4) currentLine[x + 3] = (byte)color.a; } } bmp.UnlockBits(bitmapData); }
2. 消除循环内的对象创建
循环内反复创建rgba和xy对象会触发频繁GC,拖慢运行速度。建议:
- 将
rgba和xy定义为struct(栈分配,无GC开销) - 提前声明对象,循环内复用赋值
{ BitmapData bitmapData = bmp.LockBits(new Rectangle(0, 0, bmp.Width, bmp.Height), ImageLockMode.ReadWrite, bmp.PixelFormat); int bytesPerPixel = Image.GetPixelFormatSize(bmp.PixelFormat) / 8; int heightInPixels = bitmapData.Height; int widthInBytes = bitmapData.Width * bytesPerPixel; byte* PtrFirstPixel = (byte*)bitmapData.Scan0; // 提前声明结构体,循环内复用 rgba color = default; xy pixelPos = default; for(int y = 0; y < heightInPixels; y++) { byte* currentLine = PtrFirstPixel + (y * bitmapData.Stride); pixelPos.y = y; for (int x = 0; x < widthInBytes; x += bytesPerPixel) { // 直接赋值,避免创建新对象 color.r = currentLine[x]; color.g = currentLine[x + 1]; color.b = currentLine[x + 2]; color.a = bytesPerPixel == 4 ? currentLine[x + 3] : (byte)-1; pixelPos.x = x / bytesPerPixel; for (int i = 0; i < Shaders.Count; i++) { Shaders[i].pixel(pixelPos, color); } currentLine[x] = (byte)color.r; currentLine[x + 1] = (byte)color.g; currentLine[x + 2] = (byte)color.b; if (bytesPerPixel == 4) currentLine[x + 3] = (byte)color.a; } } bmp.UnlockBits(bitmapData); }
3. 并行化处理行数据
利用多核CPU并行处理每一行(各行无依赖),用Parallel.For替代普通for循环:
{ BitmapData bitmapData = bmp.LockBits(new Rectangle(0, 0, bmp.Width, bmp.Height), ImageLockMode.ReadWrite, bmp.PixelFormat); int bytesPerPixel = Image.GetPixelFormatSize(bmp.PixelFormat) / 8; int heightInPixels = bitmapData.Height; int widthInBytes = bitmapData.Width * bytesPerPixel; byte* PtrFirstPixel = (byte*)bitmapData.Scan0; int stride = bitmapData.Stride; // 并行处理每一行 Parallel.For(0, heightInPixels, y => { byte* currentLine = PtrFirstPixel + (y * stride); rgba color = default; xy pixelPos = default; pixelPos.y = y; for (int x = 0; x < widthInBytes; x += bytesPerPixel) { color.r = currentLine[x]; color.g = currentLine[x + 1]; color.b = currentLine[x + 2]; color.a = bytesPerPixel == 4 ? currentLine[x + 3] : (byte)-1; pixelPos.x = x / bytesPerPixel; foreach (var shader in Shaders) { shader.pixel(pixelPos, color); } currentLine[x] = (byte)color.r; currentLine[x + 1] = (byte)color.g; currentLine[x + 2] = (byte)color.b; if (bytesPerPixel == 4) currentLine[x + 3] = (byte)color.a; } }); bmp.UnlockBits(bitmapData); }
注意:确保
Shaders的pixel方法是线程安全的,否则需要同步处理(但会抵消并行收益,优先保证Shader线程安全)。
4. 移除循环内的重复条件判断
把bytesPerPixel == 4的判断提到循环外,分分支处理,避免循环内的分支预测失败:
{ BitmapData bitmapData = bmp.LockBits(new Rectangle(0, 0, bmp.Width, bmp.Height), ImageLockMode.ReadWrite, bmp.PixelFormat); int bytesPerPixel = Image.GetPixelFormatSize(bmp.PixelFormat) / 8; int heightInPixels = bitmapData.Height; int widthInBytes = bitmapData.Width * bytesPerPixel; byte* PtrFirstPixel = (byte*)bitmapData.Scan0; int stride = bitmapData.Stride; if (bytesPerPixel == 4) { Parallel.For(0, heightInPixels, y => { byte* currentLine = PtrFirstPixel + (y * stride); rgba color = default; xy pixelPos = default; pixelPos.y = y; for (int x = 0; x < widthInBytes; x += 4) { color.r = currentLine[x]; color.g = currentLine[x+1]; color.b = currentLine[x+2]; color.a = currentLine[x+3]; pixelPos.x = x / 4; foreach (var shader in Shaders) { shader.pixel(pixelPos, color); } currentLine[x] = (byte)color.r; currentLine[x+1] = (byte)color.g; currentLine[x+2] = (byte)color.b; currentLine[x+3] = (byte)color.a; } }); } else // 处理3字节格式 { Parallel.For(0, heightInPixels, y => { byte* currentLine = PtrFirstPixel + (y * stride); rgba color = default; xy pixelPos = default; pixelPos.y = y; for (int x = 0; x < widthInBytes; x += 3) { color.r = currentLine[x]; color.g = currentLine[x+1]; color.b = currentLine[x+2]; color.a = (byte)-1; pixelPos.x = x / 3; foreach (var shader in Shaders) { shader.pixel(pixelPos, color); } currentLine[x] = (byte)color.r; currentLine[x+1] = (byte)color.g; currentLine[x+2] = (byte)color.b; } }); } bmp.UnlockBits(bitmapData); }
5. 可选:使用更高效的像素格式
优先选择32位ARGB格式(4字节对齐),比24位格式的内存访问速度更快,还能简化代码(无需判断字节数)。
内容的提问来源于stack exchange,提问作者Flexan
相关产品推荐
相关产品推荐

