如何逐行而非逐像素写入缓冲区,优化WebM帧的缓冲区写入?
Your current pixel-by-pixel loop gets the job done, but it’s inefficient—manual per-byte operations miss out on the low-level optimizations built into standard library functions. Let’s refactor this to use row-wise copies and add other tweaks to speed things up.
Core Optimization: Replace Pixel Loop with memcpy
The biggest performance win comes from ditching the inner pixel loop and using memcpy to copy entire rows at once. memcpy is heavily optimized by compilers (it often leverages SIMD instructions or bulk memory operations under the hood) and will drastically outperform a manual byte-by-byte loop, especially for high-resolution frames.
Here’s the refactored code:
while (videoDec->getImage(*image) == VPXDecoder::NO_ERROR) { const int w = image->getWidth(p); const int h = image->getHeight(p); // Ensure framesArray[count] has pre-allocated enough memory (w * h bytes) unsigned char* targetBuffer = framesArray[count]; // Use [] instead of at() for speed int offset = 0; for (int y = 0; y < h; y++) { // Calculate the start of the current row in the target buffer unsigned char* targetRow = targetBuffer + (w * y); // Copy the entire row from the decoder's plane to the target buffer memcpy(targetRow, image->planes[p] + offset, w); // Move to the next row in the decoder's plane offset += image->linesize[p]; } count++; // Don't forget to increment your frame counter! }
Key Changes Explained:
memcpyinstead of pixel loop: Copieswbytes (one full row) in a single call, leveraging optimized system-level memory operations.framesArray[count]instead ofat():at()performs bounds checking which adds overhead—if you’ve already ensuredcountis within valid indices (e.g., by preallocatingframesArrayto the correct size),[]is faster.- Explicit target row calculation: Makes the code clearer and avoids redundant math inside the loop.
Additional Optimizations to Consider
Preallocate
framesArrayupfront:
Before processing frames, calculate the total number of frames you expect and preallocate all buffers at once. This avoids repeated memory allocation overhead during decoding:// Example: Preallocate 100 frames of size w x h const int expectedFrames = 100; framesArray.resize(expectedFrames); for (int i = 0; i < expectedFrames; i++) { framesArray[i] = new unsigned char[w * h]; // Or use std::vector<unsigned char> for RAII safety }Pro tip: Use
std::vector<std::vector<unsigned char>>instead ofvector<unsigned char*>to avoid manual memory management and ensure safe cleanup.Check for aligned memory:
Some decoders output frames to memory aligned to 16/32-byte boundaries for SIMD compatibility. If your target buffer is also aligned,memcpy(or specialized functions likememcpy_sse2) can run even faster. Most modern allocators (likeneworstd::allocator) handle alignment automatically, but it’s worth verifying if you’re pushing for maximum performance.Avoid unnecessary copies (if possible):
Check if your VPX decoder API allows you to pass your own buffer togetImage()directly. If you can have the decoder write straight intoframesArray[count], you eliminate the copy step entirely—this is the most efficient approach if supported.
Why This Works
Manual pixel loops force the CPU to handle each byte individually, which is slow for large frames. memcpy uses optimized assembly routines that move chunks of memory at once, taking advantage of CPU cache and vectorized instructions. For a 1920x1080 frame, this means reducing 2,073,600 individual operations to just 1080 calls to memcpy—a massive performance boost.
内容的提问来源于stack exchange,提问作者Silviu Petrut

