PREEMPT_RT内核未呈现预期抢占特性的技术问题咨询
Let's break down your problem clearly: this is NOT expected behavior for a properly configured PREEMPT_RT system, and you're hitting a common pitfall with how timerfd interacts with real-time scheduling priorities.
Root Cause Analysis
Your issue stems from a mismatch between the priority of your user-space real-time threads and the kernel thread responsible for delivering timerfd events:
- Timerfd event delivery depends on
ktimersoftd: When you usetimerfd_create(CLOCK_REALTIME, 0), the timer expiration event is processed by the kernel threadktimersoftd, which runs at SCHED_FIFO priority 1 by default. - Your high-priority threads are starving
ktimersoftd: Thread 2 runs at SCHED_FIFO 54, which is far higher thanktimersoftd's priority. When Thread 2 is executing its 150ms computation, it completely preemptsktimersoftd—the kernel thread can't run at all until Thread 2 finishes. - Timerfd read blocks until
ktimersoftdruns: Thread 1'sread()on the timerfd won't return untilktimersoftdhas processed the timer expiration and queued the event. Sincektimersoftdis stuck waiting for Thread 2, Thread 1's timer logic gets delayed by the full duration of Thread 2's computation.
Your cyclictest run didn't catch this because cyclictest measures scheduling latency directly using high-priority threads that don't rely on low-priority kernel threads like ktimersoftd for event delivery.
Where You Might Have Misconfigured/Misunderstood
- Assuming timerfd events inherit user thread priority: You set Thread 1 to priority 55, but the actual event delivery depends on the kernel thread's priority, not your user thread's.
- Using one-shot timers with manual rearm: Your code uses
itime.it_interval = 0, forcing you to rearm the timer after each read. This adds unnecessary overhead and creates a window where a delayedktimersoftdcan cause missed or lagged timers. - RT bandwidth control (in 5.10 kernel): When you moved to 5.10, the kernel's RT scheduler throttling kicked in. By default, PREEMPT_RT limits the total CPU time real-time threads can use per period (default: 950ms runtime every 1s). If Thread 2's 150ms computation runs too frequently, it exceeds this limit, and the scheduler throttles all RT threads—including your timer thread.
Fixes and Workarounds
Here are actionable solutions to resolve the issue, ordered by effectiveness:
1. Boost the ktimersoftd thread priority
The simplest fix is to raise the priority of ktimersoftd above all your user-space real-time threads. This ensures it can preempt any user thread to deliver timer events:
# Find the PID of ktimersoftd KTIMER_PID=$(pgrep ktimersoftd) # Set it to SCHED_FIFO priority 56 (higher than your Thread 2's 54) chrt -f -p 56 $KTIMER_PID
To make this persistent across reboots, add this command to your system's startup scripts.
2. Use periodic timers instead of one-shot
Instead of rearming the timer manually after each read, configure it as a periodic timer. This reduces system call overhead and lets the kernel manage timer scheduling more efficiently:
// In rearmTimer, set the interval to match the value itime.it_interval.tv_sec = ts.tv_sec; itime.it_interval.tv_nsec = ts.tv_nsec;
This way, you only need to call timerfd_settime once at startup, not after every read.
3. Use CLOCK_MONOTONIC instead of CLOCK_REALTIME
CLOCK_REALTIME can be adjusted (e.g., via NTP), which can introduce unexpected timer behavior. CLOCK_MONOTONIC is a steady, non-adjustable clock ideal for periodic tasks:
g_timer = timerfd_create(CLOCK_MONOTONIC, 0);
4. Adjust RT scheduler bandwidth control (for 5.10+)
If you hit throttling on 5.10, you can either:
- Increase the allowed runtime for RT threads:
echo 990000 > /proc/sys/kernel/sched_rt_runtime_us echo 1000000 > /proc/sys/kernel/sched_rt_period_us - Or disable bandwidth control entirely (use with caution—this can let RT threads starve non-RT threads):
echo -1 > /proc/sys/kernel/sched_rt_runtime_us
5. Switch to pthread_timer for user-space event delivery
For more control, use pthread_timer_create instead of timerfd. This lets you attach a callback to your timer that runs directly in your high-priority Thread 1's context, bypassing ktimersoftd entirely. The callback will run at Thread 1's priority, so it won't be preempted by Thread 2.
Verification
After applying any of these fixes, run your test application again. The timer interval should stay consistent within a few microseconds (expected for PREEMPT_RT), even while Thread 2 is running its computation.
内容的提问来源于stack exchange,提问作者BoozyWoozy

