如何在C++中实现±150纳秒精度的指令间隔控制并优化GPIO代码
树莓派GPIO精确计时优化问题
我先编写了Java程序并转为C版本,但C程序运行未达预期(比预想的慢),推测问题出在操作间的时间控制方式上。日常开发以Java为主,测试发现Java操作树莓派3B的GPIO输出比C慢40-65倍,但当前C代码的性能表现不佳。树莓派4B上单次digitalWrite调用耗时:Java通过JNI为3500-6500纳秒,C仅88纳秒。现需优化C代码,实现digitalWrite操作间的精确计时,精度要求±150纳秒。
我的Java代码(树莓派4B上JNI单次调用digitalWrite(int,int)耗时3500-6500纳秒)
public static void switchOnLeds(int firstLed, int lastLed){ final int BIT_0_HIGH_VOLTAGE_TIME = 250; // ws2811快速模式下逻辑0高电平持续时间(纳秒) final int BIT_0_LOW_VOLTAGE_TIME = 1000; // ws2811快速模式下逻辑0低电平持续时间(纳秒) final int BIT_1_HIGH_VOLTAGE_TIME = 600; // ws2811快速模式下逻辑1高电平持续时间(纳秒) final int BIT_1_LOW_VOLTAGE_TIME = 650; // ws2811快速模式下逻辑1低电平持续时间(纳秒) final int BITS_FOR_SINGLE_LED = 24; // 单颗LED对应的信号位数 final int PIN = 4; // GPIO输出引脚 final int LEDS = 10; long start, end; for (int i = 0; i < LEDS; i++){ if (i >= firstLed && i <= lastLed){ for (int bit = 0; bit < BITS_FOR_SINGLE_LED; bit++){ start = System.nanoTime(); end = start + BIT_1_HIGH_VOLTAGE_TIME; digitalWrite(PIN, true); while (System.nanoTime() < end){ // 等待 } start = System.nanoTime(); end = start + BIT_1_LOW_VOLTAGE_TIME; digitalWrite(PIN, false); while (System.nanoTime() < end){ // 等待 } } } else { for (int bit = 0; bit < BITS_FOR_SINGLE_LED; bit++){ start = System.nanoTime(); end = start + BIT_0_HIGH_VOLTAGE_TIME; digitalWrite(PIN, true); while (System.nanoTime() < end){ // 等待 } start = System.nanoTime(); end = start + BIT_0_LOW_VOLTAGE_TIME; digitalWrite(PIN, false); while (System.nanoTime() < end){ // 等待 } } } } }
待优化的C++代码(树莓派4B上单次调用digitalWrite(int,int)耗时88纳秒)
void switchOnLeds(int firstLed, int lastLed) { const int BIT_0_HIGH_VOLTAGE_TIME = 250; // ws2811快速模式下逻辑0高电平持续时间(纳秒) const int BIT_0_LOW_VOLTAGE_TIME = 1000; // ws2811快速模式下逻辑0低电平持续时间(纳秒) const int BIT_1_HIGH_VOLTAGE_TIME = 600; // ws2811快速模式下逻辑1高电平持续时间(纳秒) const int BIT_1_LOW_VOLTAGE_TIME = 650; // ws2811快速模式下逻辑1低电平持续时间(纳秒) const int BITS_FOR_SINGLE_LED = 24; // 单颗LED对应的信号位数 const int PIN = 4; // GPIO输出引脚 const int LEDS = 10; std::chrono::steady_clock::time_point start = std::chrono::high_resolution_clock::now(); std::chrono::steady_clock::time_point end = std::chrono::high_resolution_clock::now(); for (int i = 0; i < LEDS; i++) { if (i >= firstLed && i <= lastLed) { for (int bit = 0; bit < BITS_FOR_SINGLE_LED; bit++) { start = std::chrono::high_resolution_clock::now(); digitalWrite(PIN, true); while ((std::chrono::high_resolution_clock::now() - start).count() < BIT_1_HIGH_VOLTAGE_TIME) { // 等待 } start = std::chrono::high_resolution_clock::now(); digitalWrite(PIN, false); while ((std::chrono::high_resolution_clock::now() - start).count() < BIT_1_LOW_VOLTAGE_TIME) { // 等待 } } } else { for (int bit = 0; bit < BITS_FOR_SINGLE_LED; bit++) { start = std::chrono::high_resolution_clock::now(); digitalWrite(PIN, true); while ((std::chrono::high_resolution_clock::now() - start).count() < BIT_0_HIGH_VOLTAGE_TIME) { // 等待 } start = std::chrono::high_resolution_clock::now(); digitalWrite(PIN, false); while ((std::chrono::high_resolution_clock::now() - start).count() < BIT_0_LOW_VOLTAGE_TIME) { // 等待 } } } } }
优化方案
1. 使用硬件寄存器级计时,替代std::chrono
树莓派4B的ARM架构支持CNTVCT_EL0系统计数器寄存器,直接读取它能获得几纳秒级的计时开销,远低于std::chrono的调用成本:
#include <cstdint> // 树莓派4B系统时钟为24MHz,计数器每tick对应1/24e6秒,转换为纳秒需乘以(1000000000 / 24) inline uint64_t get_ns() { uint64_t val; asm volatile("mrs %0, cntvct_el0" : "=r"(val)); return val * (1000000000 / 24); }
2. 预计算结束时间,减少循环内运算
参考Java代码逻辑,提前计算结束时间,避免循环内重复计算时间差:
// 以逻辑1的高电平等待为例 uint64_t start = get_ns(); uint64_t end = start + BIT_1_HIGH_VOLTAGE_TIME; digitalWrite(PIN, true); while (get_ns() < end) { // 加入nop指令防止编译器优化空循环 asm volatile("nop"); }
3. 阻止编译器优化空循环
空循环会被编译器优化为无操作,导致等待时间失效。加入asm volatile("nop")或空汇编指令,强制编译器保留循环:
while (get_ns() < end) { asm volatile(""); // 阻止优化 }
4. 提取重复逻辑,减少分支判断
将BIT0和BIT1的发送逻辑封装为函数,减少主循环内的分支判断,提升执行效率:
void send_bit(bool is_high, int PIN) { const int high_time = is_high ? BIT_1_HIGH_VOLTAGE_TIME : BIT_0_HIGH_VOLTAGE_TIME; const int low_time = is_high ? BIT_1_LOW_VOLTAGE_TIME : BIT_0_LOW_VOLTAGE_TIME; uint64_t start = get_ns(); uint64_t end = start + high_time; digitalWrite(PIN, true); while (get_ns() < end) { asm volatile("nop"); } start = get_ns(); end = start + low_time; digitalWrite(PIN, false); while (get_ns() < end) { asm volatile("nop"); } }
主循环中直接调用该函数即可:
for (int i = 0; i < LEDS; i++) { bool is_high_led = (i >= firstLed && i <= lastLed); for (int bit = 0; bit < BITS_FOR_SINGLE_LED; bit++) { send_bit(is_high_led, PIN); } }
5. 直接操作GPIO寄存器,跳过库函数封装
如果使用wiringPi等库的digitalWrite,可以直接映射GPIO内存并操作寄存器,进一步降低GPIO操作耗时:
#include <sys/mman.h> #include <fcntl.h> #include <unistd.h> #define GPIO_BASE 0xFE200000 // 树莓派4B的GPIO基地址 volatile uint32_t* gpio = nullptr; // 初始化GPIO内存映射 void init_gpio() { int mem_fd = open("/dev/mem", O_RDWR | O_SYNC); gpio = (volatile uint32_t*)mmap(nullptr, 4096, PROT_READ | PROT_WRITE, MAP_SHARED | MAP_PHYS, mem_fd, GPIO_BASE); close(mem_fd); // 设置PIN为输出模式(以PIN4为例) int pin = 4; int reg = pin / 10; int shift = (pin % 10) * 3; gpio[reg] &= ~(7 << shift); gpio[reg] |= (1 << shift); // 设置为输出 } // 快速GPIO写操作 inline void fast_digitalWrite(int pin, bool high) { if (high) { gpio[7] = 1 << pin; // SET寄存器,置高电平 } else { gpio[10] = 1 << pin; // CLEAR寄存器,置低电平 } }
使用fast_digitalWrite替代原digitalWrite,耗时可降至几纳秒级别。
内容的提问来源于stack exchange,提问作者Alexander Gorodilov
相关产品推荐
相关产品推荐

