C++ AVX2图像提亮函数通过C# P/Invoke调用时性能大幅下降的原因及优化方案咨询
C++ AVX2图像提亮函数通过C# P/Invoke调用时性能大幅下降的原因及优化方案咨询
我有一个使用AVX2 intrinsics实现的C图像提亮函数,直接在C中测试时,处理3840×240分辨率的图像大约需要500微秒,但通过C#的P/Invoke调用同一个函数时,耗时却达到了约4毫秒,性能差距非常明显。
不同分辨率下的测试数据如下:
- 3840×2160图像:原生C++耗时约1.5ms,P/Invoke调用耗时约4.5ms
- 10000×4000图像:原生C++耗时约3ms,P/Invoke调用耗时约8ms
我的实现环境是:
- C++:原生函数使用AVX2指令处理图像数据
- C#:通过P/Invoke调用该函数,直接按引用传递图像数据
C# 主程序
static void Main(string[] args) { int width = 3840; int height = 240; byte[] image = new byte[width * height]; Random random = new Random(); // Fill the image array with random brightness values between 0 and 255 for (int i = 0; i < image.Length; i++) { image[i] = (byte)random.Next(0, 256); } byte brightness = 30; // Measure C# processing time Stopwatch sw = Stopwatch.StartNew(); ProcessorCSharp.BrightenImage(image, brightness); sw.Stop(); Console.WriteLine("C# Time: {0} microseconds", sw.Elapsed.TotalMilliseconds * 1000); }
C# 包装类
public class ProcessorCSharp { [DllImport("ImageProcessingLib.dll", CallingConvention = CallingConvention.Cdecl)] private static extern void brightenImageSIMD(IntPtr image, int size, byte brightness); public static unsafe void BrightenImage(byte[] image, byte brightness) { int size = image.Length; fixed (byte* p = image) { brightenImageSIMD((IntPtr)p, size, brightness); } } }
C++ 函数
#include <immintrin.h> // AVX2 intrinsics #include <vector> #include <algorithm> // For std::min extern "C" __declspec(dllexport) // in my main c++ code this line doesn't exist void brightenImageSIMD(uint8_t* image, size_t size, uint8_t brightness) { size_t i = 0; __m256i brightnessVector = _mm256_set1_epi8(brightness); __m256i maxVector = _mm256_set1_epi8(255); for (; i + 31 < size; i += 32) { __m256i pixels = _mm256_loadu_si256((__m256i*) &image[i]); __m256i brightened = _mm256_adds_epu8(pixels, brightnessVector); __m256i clamped = _mm256_min_epu8(brightened, maxVector); _mm256_storeu_si256((__m256i*) &image[i], clamped); } for (; i < size; ++i) { image[i] = std::min(image[i] + brightness, 255); } } // Helper function to measure execution time template <typename Func, typename... Args> long long measureExecutionTime(Func func, Args&&... args) { auto start = std::chrono::high_resolution_clock::now(); func(std::forward<Args>(args)...); auto end = std::chrono::high_resolution_clock::now(); return std::chrono::duration_cast<std::chrono::microseconds>(end - start).count(); } int main() { const int width = 3840; const int height = 2160; const uint8_t brightnessIncrease = 30; std::vector<uint8_t> image(width * height); // Set up random number generation std::random_device rd; // Seed for the random number engine std::mt19937 gen(rd()); // Standard mersenne_twister_engine std::uniform_int_distribution<> dis(0, 255); // Range from 0 to 255 for 8-bit brightness levels // Fill the image with random values for (auto& pixel : image) { pixel = static_cast<uint8_t>(dis(gen)); } // Measure performance of SIMD method auto imageCopy3 = image; // Make another copy for fair comparison long long timeSIMD = measureExecutionTime(brightenImageSIMD, imageCopy3, brightnessIncrease); std::cout << "SIMD (AVX2) method time: " << timeSIMD << " microseconds" << std::endl; return 0; }
核心问题
我尝试用C包装函数来提升C#实时图像处理的性能,但实际调用时的性能表现却远不如原生C。我已经使用了unsafe代码和fixed指针来避免数组拷贝,但性能差距依然很大。
我的疑问
- 为什么原生C++执行和C# P/Invoke调用之间存在这么大的性能差距?
- 我能做些什么让C#调用版本的性能接近原生C++?
- 像OpenCvSharp这样的库通过P/Invoke调用原生OpenCV函数依然能保持高性能,它们用到了哪些我没注意到的技术?
补充说明:我使用的是.NET Framework 4.8;我曾将函数运行100次取平均值,此时两者的耗时差距缩小了,但我的使用场景中无法通过多次调用让函数速度变快。
备注:内容来源于stack exchange,提问作者MustafaVisys
相关产品推荐
相关产品推荐

