多线程系统中pthread_cond_timedwait与pthread_mutex_lock崩溃问题排查
Alright, let's break down what's happening here and fix this crash step by step.
From your code snippet and crash details, the main issues stem from incorrect condition variable usage, flawed parameter passing, and risky timeout handling:
1. You're passing the Data struct by value (fatal mistake)
Your Thread A function signature is void sleep(Data x)—this passes the entire Data struct as a copy. That means Thread A is operating on a local stack copy of the struct, not the same shared instance Thread B is using.
This creates a disjoint state:
- Thread A locks a mutex that only exists in its own stack frame
- When Thread A exits
sleep, that copy (and its mutex/condition variable) is destroyed - Thread B ends up waiting on a completely different mutex, or worse—if it's referencing a destroyed mutex, that directly causes crashes. The stack trace showing
value = 0confirms this: Thread B is looking at a differentvaluethan the one Thread A modified.
2. Condition variables aren't paired with proper condition checks
POSIX condition variables have a hard rule: you must wrap pthread_cond_wait/pthread_cond_timedwait in a loop that checks your actual wait condition. This handles spurious wakeups (wait can return without a signal) and ensures you don't act on invalid state after a timeout.
In Thread A, you set value = 0 immediately after the wait finishes—no matter if it was woken up normally or timed out. This clobbers the value = 1 state Thread B sets, creating a race that breaks your synchronization logic.
3. Global timeout variable causes race hazards
It looks like T (the timespec for the timeout) is a global variable. If multiple threads call sleep, or even if Thread A calls it multiple times, they'll overwrite each other's timeout values. This makes the timeout behavior unpredictable, amplifying the race conditions that trigger your crash.
Let's address these issues one by one:
1. Pass Data by pointer to share state
First, update both functions to accept a pointer to Data—this ensures both threads operate on the exact same shared struct:
void sleep(Data *x) { // All operations now use the shared struct via pointer } void wakeup(Data *y) { // Same here—operate on the shared instance }
2. Wrap condition wait in a loop with proper checks
Modify Thread A to only wait while the condition isn't met (x->value != 1), and only reset value if the wait was successfully woken up (not timed out):
void sleep(Data *x) { pthread_mutex_lock( &(x->lock) ); // Declare timeout as a local variable to avoid race conditions struct timespec T; struct timeval _T; gettimeofday( &_T, NULL ); // Set your desired timeout here (example: 1 second from now) T.tv_sec = _T.tv_sec + 1; T.tv_nsec = _T.tv_usec * 1000; // Convert microseconds to nanoseconds // Loop to handle spurious wakeups and timeouts while (x->value != 1) { int ret = pthread_cond_timedwait( &(x->condition), &(x->lock), &T ); if (ret == ETIMEDOUT) { // Handle timeout (e.g., log it) — don't force value to 0! break; } } // Only reset value if we were actually woken up by Thread B if (x->value == 1) { x->value = 0; } pthread_mutex_unlock( &(x->lock) ); }
3. Keep Thread B's logic (with pointer updates)
Thread B's core logic is correct—just ensure it uses the shared struct pointer:
void wakeup(Data *y) { pthread_mutex_lock( &(y->lock) ); y->value = 1; pthread_cond_signal( &(y->condition) ); pthread_mutex_unlock( &(y->lock) ); }
- Shared state via pointer: Both threads now use the same mutex, condition variable, and
value—no more disjoint synchronization primitives. - Condition loop: The loop ensures we only proceed when the expected state is met, avoiding spurious wakeups and incorrect state modifications after timeouts.
- Local timeout variable: Eliminates race conditions where multiple threads overwrite the timeout value.
- Targeted value reset: We only reset
valueto 0 when we know Thread B successfully woke us up, so Thread B's state changes aren't clobbered unexpectedly.
内容的提问来源于stack exchange,提问作者Anu

