嵌入式系统下C语言无循环实现跳固定字节的内存拷贝
Great question—when dealing with large embedded data transfers, ditching slow byte-by-byte loops for faster, hardware-assisted or vectorized operations is key. Let’s break down the best approaches for your use case (skipping 1 byte between each copied element, scalable to datasets like 1500 bytes):
If your MCU supports memory-to-memory DMA transfers (most modern ones do), this is hands down the fastest solution. DMA handles the transfer entirely in hardware without CPU intervention, so you’ll get near-zero overhead even for large datasets.
Here’s how to configure it for your example:
- Set the DMA channel to memory-to-memory mode
- Configure the source address increment to
2(to skip 1 byte after each copied element) - Set the destination address increment to
1(normal sequential writes) - Set the transfer length to
750(for 1500 source bytes, since you take every other byte) - Start the transfer and either poll the completion flag or use an interrupt to know when it’s done.
No loops, no wasted CPU cycles—perfect for real-time embedded systems.
If DMA isn’t an option (e.g., your hardware lacks it, or you need more flexibility), use SIMD (Single Instruction, Multiple Data) intrinsics. These let you process multiple bytes in parallel with a single CPU instruction, drastically cutting down on processing time.
For ARM-based systems with NEON support (common in many MCUs), here’s a simplified snippet to copy every other byte:
#include <arm_neon.h> void skip_copy_neon(const uint8_t *src, uint8_t *dst, size_t num_dst_elements) { size_t i; // Process 8 source bytes (4 destination elements) in parallel for (i = 0; i < num_dst_elements - 3; i += 4) { uint8x8_t src_vec = vld1_u8(src + 2*i); // Load 8 consecutive source bytes // Extract even-indexed bytes from the vector uint8x4_t dst_vec = vget_low_u8(vzip1_u8(src_vec, src_vec)); vst1_u8(dst + i, dst_vec); // Store 4 bytes to destination } // Handle remaining elements with a tiny, negligible loop for (; i < num_dst_elements; i++) { dst[i] = src[2*i]; } }
While this has a loop, it processes 4 elements at a time instead of 1—for 1500 bytes, that’s only 188 iterations instead of 750. Most compilers can also auto-vectorize simple loops if you enable optimizations (e.g., -O3 for GCC), but using intrinsics gives you direct control.
On x86 systems, you’d use SSE/AVX intrinsics like _mm_loadu_si128 and _mm_shuffle_epi8 to extract the desired bytes.
Before rolling your own SIMD code, check if your compiler can optimize your existing loop automatically. Enable high optimization levels and add hints to help the compiler vectorize the code:
#pragma GCC optimize("O3,unroll-loops") void skip_copy_optimized(const uint8_t *src, uint8_t *dst, size_t num_dst_elements) { // Tell the compiler arrays are aligned (critical for vectorization) __builtin_assume_aligned(src, 16); __builtin_assume_aligned(dst, 16); // Hint that src and dst don't overlap __builtin_memcpy(dst, src, 0); // Dummy call to enforce non-overlap for (size_t i = 0; i < num_dst_elements; i++) { dst[i] = src[2*i]; } }
With -O3, GCC will often unroll this loop and generate NEON/SSE instructions automatically, cutting your 5-6ms runtime down to well under 1ms.
If you’re dealing with a fixed, small dataset (like your 6-byte example), you can avoid loops entirely using struct aliasing:
typedef struct { uint8_t val0; uint8_t skip0; uint8_t val1; uint8_t skip1; uint8_t val2; uint8_t skip2; } __attribute__((packed)) SkipStruct; const SkipStruct *src_struct = (const SkipStruct *)source; uint8_t destination[3] = {src_struct->val0, src_struct->val1, src_struct->val2};
This works for fixed small sizes but isn’t scalable to 1500 bytes—think of it as a trick for tiny, known-length transfers.
Key Takeaways
- Always ensure data alignment when using DMA or SIMD—misaligned accesses can cause crashes or slowdowns. Use
__attribute__((aligned(16)))on your arrays if needed. - DMA is the top choice for embedded systems because it offloads work from the CPU, freeing it up for other real-time tasks.
- If you must use software-based copies, SIMD intrinsics or compiler-optimized loops will drastically reduce overhead compared to naive byte-by-byte loops.
内容的提问来源于stack exchange,提问作者Ameer Hamza

